VLDeformer: Vision-Language Decomposed Transformer for Fast Cross-Modal Retrieval

Lisai Zhang; Hongfa Wu; Qingcai Chen; Yimeng Deng; Zhonghua Li; Dejiang Kong; Zhao Cao; Joanna Siebert; Yunpeng Han

VLDeformer：高速クロスモーダル検索のための視覚言語分解トランスフォーマー

クロスモデル検索は、テキストのみの検索エンジン（SE）の最も重要なアップグレードの1つとして浮上しています。最近、初期の相互作用を介したペアワイズテキスト画像入力の強力な表現により、視覚言語（VL）トランスフォーマーの精度は、テキスト画像検索の既存の方法を上回っています。ただし、同じパラダイムを推論に使用した場合、VLトランスの効率はまだ低すぎて、実際のクロスモーダルSEに適用できません。人間の学習のメカニズムとクロスモーダル知識の使用に触発されて、この論文は、VLトランスフォーマーの優れた精度を維持しながら効率を大幅に向上させる新しい視覚言語分解トランスフォーマー（VLDeformer）を紹介します。提案手法により、モデル間検索は、VL変換器学習段階とVL分解段階の2つの段階に分けられます。後の段階は、単一のモーダルインデックスの役割を果たします。これは、テキストSEのインデックスという用語にある程度似ています。モデルは、初期の相互作用の事前トレーニングからクロスモーダルの知識を学習し、個々のエンコーダーに分解されます。分解は、監視のために小さなターゲットデータセットのみを必要とし、1000倍以上の加速と0.6％未満の平均リコールドロップの両方を達成します。 VLDeformerは、COCOおよびFlickr30kの最先端の視覚的セマンティック埋め込み方法よりも優れています。

Cross-model retrieval has emerged as one of the most important upgrades for text-only search engines (SE). Recently, with powerful representation for pairwise text-image inputs via early interaction, the accuracy of vision-language (VL) transformers has outperformed existing methods for text-image retrieval. However, when the same paradigm is used for inference, the efficiency of the VL transformers is still too low to be applied in a real cross-modal SE. Inspired by the mechanism of human learning and using cross-modal knowledge, this paper presents a novel Vision-Language Decomposed Transformer (VLDeformer), which greatly increases the efficiency of VL transformers while maintaining their outstanding accuracy. By the proposed method, the cross-model retrieval is separated into two stages: the VL transformer learning stage, and the VL decomposition stage. The latter stage plays the role of single modal indexing, which is to some extent like the term indexing of a text SE. The model learns cross-modal knowledge from early-interaction pre-training and is then decomposed into an individual encoder. The decomposition requires only small target datasets for supervision and achieves both 1000+ times acceleration and less than 0.6% average recall drop. VLDeformer also outperforms state-of-the-art visual-semantic embedding methods on COCO and Flickr30k.

updated: Mon Nov 22 2021 02:58:47 GMT+0000 (UTC)

published: Wed Oct 20 2021 09:00:51 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト