Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot Learning

Xiaocheng Lu; Ziming Liu; Song Guo; Jingcai Guo

合成ゼロショット学習のための分解されたソフトプロンプトガイド付きフュージョン強化

合成ゼロショット学習 (CZSL) は、トレーニング中に既知の状態とオブジェクトによって形成された新しい概念を認識することを目的としています。既存の方法は、結合された状態とオブジェクトの表現を学習して、目に見えない構成の一般化に挑戦するか、2 つの分類子を設計して、状態とオブジェクトを画像の特徴から別々に識別し、それらの間の固有の関係を無視します。上記の問題を共同で排除し、より堅牢な CZSL システムを構築するために、目に見えない構成認識に視覚言語モデル (VLM) を含めることにより、Decomposed Fusion with Soft Prompt (DFSP)1 と呼ばれる新しいフレームワークを提案します。具体的には、DFSP は学習可能なソフトプロンプトと状態およびオブジェクトのベクトルの組み合わせを構築し、それらの結合表現を確立します。さらに、クロスモーダル分解融合モジュールは、言語と画像の枝の間で設計されており、画像の特徴の代わりに言語の特徴の間で状態とオブジェクトを分解します。特に、分解された特徴と融合することで、画像の特徴は、それぞれ状態とオブジェクトとの関係を学習するためにより表現力豊かになり、ペア空間での目に見えない構成の応答を改善し、見られるセットと見られないセットの間のドメインギャップを狭めます。 3 つの挑戦的なベンチマークでの実験結果は、私たちのアプローチが他の最先端の方法よりも大幅に優れていることを示しています。

Compositional Zero-Shot Learning (CZSL) aims to recognize novel concepts formed by known states and objects during training. Existing methods either learn the combined state-object representation, challenging the generalization of unseen compositions, or design two classifiers to identify state and object separately from image features, ignoring the intrinsic relationship between them. To jointly eliminate the above issues and construct a more robust CZSL system, we propose a novel framework termed Decomposed Fusion with Soft Prompt (DFSP)1, by involving vision-language models (VLMs) for unseen composition recognition. Specifically, DFSP constructs a vector combination of learnable soft prompts with state and object to establish the joint representation of them. In addition, a cross-modal decomposed fusion module is designed between the language and image branches, which decomposes state and object among language features instead of image features. Notably, being fused with the decomposed features, the image features can be more expressive for learning the relationship with states and objects, respectively, to improve the response of unseen compositions in the pair space, hence narrowing the domain gap between seen and unseen sets. Experimental results on three challenging benchmarks demonstrate that our approach significantly outperforms other state-of-the-art methods by large margins.

updated: Sat Nov 19 2022 12:29:12 GMT+0000 (UTC)

published: Sat Nov 19 2022 12:29:12 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト