DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Wei Li; Linchao Zhu; Longyin Wen; Yi Yang

DeCap: テキストのみのトレーニングによるゼロショットキャプション用の CLIP 潜在データのデコード

大規模な事前トレーニング済みマルチモーダルモデル (例: CLIP) は、多くの差別的なタスクで強力なゼロショット転送機能を示します。ゼロショット画像条件付きテキスト生成タスクへのそれらの適応は、ますます関心を集めています。既存の大規模な言語モデル（例えば、ＧＰＴ－２）を利用するか、エンドツーエンド方式でエンコーダ－デコーダネットワークを事前トレーニングするかのいずれかによる、ゼロショットキャプションへの従来技術のアプローチ。この作業では、ゼロショットキャプション用の DeCap という名前の単純なフレームワークを提案します。軽量の視覚認識言語デコーダーを紹介します。このデコーダーは、データ効率と計算効率の両方に優れています。1) トレーニングに必要なのはテキストデータのみであり、ペアデータの収集の負担が軽減されます。 2) エンドツーエンドのトレーニングは必要ありません。テキストのみのデータでトレーニングすると、デコーダーは、市販の CLIP エンコーダーから抽出されたテキスト埋め込みをプレフィックス埋め込みとして受け取ります。課題は、デコーダーがテキストコーパスでトレーニングされることですが、推論段階では、視覚的な入力に基づいてキャプションを生成する必要があります。モダリティギャップの問題は、マルチモーダルの対照的なモデルで広く観察されており、視覚的な埋め込みを接頭辞の埋め込みとして直接使用することができません。モダリティのギャップを減らすためのトレーニング不要のメカニズムを提案します。視覚的埋め込みをCLIPテキスト埋め込みスペースに投影しますが、投影された埋め込みは視覚入力の情報を保持します。射影された埋め込みをプレフィックス埋め込みとして取り、デコーダーは視覚入力に一致する高品質の説明を生成します。実験では、MSCOCO や NoCaps などの典型的な画像キャプションベンチマークで、DeCap が他のゼロショットキャプション方法やペアになっていないキャプション方法よりも優れていることが示されています。

Large-scale pre-trained multi-modal models (e.g., CLIP) demonstrate strong zero-shot transfer capability in many discriminative tasks. Their adaptation to zero-shot image-conditioned text generation tasks has drawn increasing interest. Prior arts approach to zero-shot captioning by either utilizing the existing large language models (e.g., GPT-2) or pre-training the encoder-decoder network in an end-to-end manner. In this work, we propose a simple framework, named DeCap, for zero-shot captioning. We introduce a lightweight visual-aware language decoder. This decoder is both data-efficient and computation-efficient: 1) it only requires the text data for training, easing the burden on the collection of paired data. 2) it does not require end-to-end training. When trained with text-only data, the decoder takes the text embedding extracted from the off-the-shelf CLIP encoder as a prefix embedding. The challenge is that the decoder is trained on the text corpus but at the inference stage, it needs to generate captions based on visual inputs. The modality gap issue is widely observed in multi-modal contrastive models that prevents us from directly taking the visual embedding as the prefix embedding. We propose a training-free mechanism to reduce the modality gap. We project the visual embedding into the CLIP text embedding space, while the projected embedding retains the information of the visual input. Taking the projected embedding as the prefix embedding, the decoder generates high-quality descriptions that match the visual input. The experiments show that DeCap outperforms other zero-shot captioning methods and unpaired captioning methods on the typical image captioning benchmarks, i.e., MSCOCO and NoCaps.

updated: Mon Mar 06 2023 11:02:47 GMT+0000 (UTC)

published: Mon Mar 06 2023 11:02:47 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト