I-Tuning: Tuning Frozen Language Models with Image for Lightweight Image Captioning

Ziyang Luo; Zhipeng Hu; Yadong Xi; Rongsheng Zhang; Jing Ma

I-Tuning: 軽量な画像キャプション用に画像を使用して凍結された言語モデルを調整する

画像キャプションは、画像の言語記述を生成することを目的とした従来の視覚と言語のタスクです。最近の研究では、モデルのサイズとトレーニングデータの数を拡大することに重点が置かれており、モデルのトレーニングのコストが大幅に増加しています。これらの高コストモデルとは異なり、少数のトレーニング可能なパラメーターを含む軽量の画像キャプションフレームワーク (I-Tuning) を導入します。トレーニング不可能な事前トレーニング済み言語デコーダー GPT2 とビジョンエンコーダー CLIP-ViT を接続するための新しい I-Tuning 相互注意モジュールを設計します。ほとんどのパラメーターはトレーニング中に更新する必要がないため、フレームワークは軽量で高速です。 3 つの画像キャプションベンチマークで実施された実験結果は、私たちのフレームワークが大規模なベースラインシステムと同等またはそれ以上のパフォーマンスを達成することを明らかにしています。しかし、私たちのモデルに含まれるトレーニング可能なパラメーターは最大 10 分の 1 であり、最先端のベースラインと比較して、トレーニングに必要なデータははるかに少なくなります。

Image Captioning is a traditional vision-and-language task that aims to generate the language description of an image. Recent studies focus on scaling up the model size and the number of training data, which significantly increase the cost of model training. Different to these heavy-cost models, we introduce a lightweight image captioning framework (I-Tuning), which contains a small number of trainable parameters. We design a novel I-Tuning cross-attention module to connect the non-trainable pre-trained language decoder GPT2 and vision encoder CLIP-ViT. Since most parameters are not required to be updated during training, our framework is lightweight and fast. Experimental results conducted on three image captioning benchmarks reveal that our framework achieves comparable or better performance than the large-scale baseline systems. But our models contain up to 10 times fewer trainable parameters and require much fewer data for training compared with state-of-the-art baselines.

updated: Mon Mar 13 2023 05:51:27 GMT+0000 (UTC)

published: Mon Feb 14 2022 09:36:50 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト