Contrastive Semantic Similarity Learning for Image Captioning Evaluation with Intrinsic Auto-encoder

Chao Zeng; Tiesong Zhao; Sam Kwong

固有のオートエンコーダによる画像キャプション評価のための対照的な意味的類似性学習

画像のキャプションの品質を自動的に評価することは非常に難しい場合があります。人間の言語は非常に柔軟であり、同じ意味でさまざまな表現が存在する可能性があるためです。現在のキャプションメトリックのほとんどは、候補キャプションとグラウンドトゥルースラベルセンテンスの間のトークンレベルのマッチングに依存しています。通常、文レベルの情報は無視されます。オートエンコーダメカニズムと対照表現学習の進歩に動機付けられて、画像キャプションの学習ベースのメトリックを提案します。これを本質的画像キャプション評価（I ^ 2CE）と呼びます。文レベルの表現を学習するために、3つのプログレッシブモデル構造を開発します。シングルブランチモデル、デュアルブランチモデル、トリプルブランチモデルです。私たちの経験的テストは、デュアルブランチ構造でトレーニングされたI ^ 2CEが、現代の画像キャプション評価メトリックに対する人間の判断とのより良い一貫性を達成することを示しています。さらに、いくつかの最先端の画像キャプションモデルを選択し、最新のメトリックと提案されたI ^ 2CEの両方に関するMSCOCOデータセットでのパフォーマンスをテストします。実験結果は、提案された方法が他の現代的な測定基準から生成されたスコアとうまく一致することができることを示しています。この懸念に関して、提案された測定基準は、既存のものを補完する可能性のある、キャプション間の固有の情報の新しい指標として役立つ可能性があります。

Automatically evaluating the quality of image captions can be very challenging since human language is quite flexible that there can be various expressions for the same meaning. Most of the current captioning metrics rely on token level matching between candidate caption and the ground truth label sentences. It usually neglects the sentence-level information. Motivated by the auto-encoder mechanism and contrastive representation learning advances, we propose a learning-based metric for image captioning, which we call Intrinsic Image Captioning Evaluation(I^2CE). We develop three progressive model structures to learn the sentence level representations--single branch model, dual branches model, and triple branches model. Our empirical tests show that I^2CE trained with dual branches structure achieves better consistency with human judgments to contemporary image captioning evaluation metrics. Furthermore, We select several state-of-the-art image captioning models and test their performances on the MS COCO dataset concerning both contemporary metrics and the proposed I^2CE. Experiment results show that our proposed method can align well with the scores generated from other contemporary metrics. On this concern, the proposed metric could serve as a novel indicator of the intrinsic information between captions, which may be complementary to the existing ones.

updated: Tue Jun 29 2021 12:27:05 GMT+0000 (UTC)

published: Tue Jun 29 2021 12:27:05 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト