Large-Scale Zero-Shot Image Classification from Rich and Diverse Textual Descriptions

Sebastian Bujwid; Josephine Sullivan

豊富で多様なテキスト記述からの大規模ゼロショット画像分類

ImageNetでゼロショット学習（ZSL）のクラスの豊富で多様なテキスト記述を使用することの影響を研究します。各ImageNetクラスを対応するウィキペディアの記事に一致させる新しいデータセットImageNet-Wikiを作成します。これらのウィキペディアの記事をクラスの説明として使用するだけで、以前の作品よりもはるかに高いZSLパフォーマンスが得られることを示します。このタイプの補助データを使用する単純なモデルでさえ、クラス名の単語埋め込みエンコーディングの標準機能に依存する最先端のモデルよりも優れています。これらの結果は、ZSLのテキストによる説明の有用性と重要性、およびアルゴリズムの進歩と比較した補助データ型の相対的な重要性を浮き彫りにしています。私たちの実験結果は、標準的なゼロショット学習アプローチがクラスのカテゴリ全体で一般化が不十分であることも示しています。

We study the impact of using rich and diverse textual descriptions of classes for zero-shot learning (ZSL) on ImageNet. We create a new dataset ImageNet-Wiki that matches each ImageNet class to its corresponding Wikipedia article. We show that merely employing these Wikipedia articles as class descriptions yields much higher ZSL performance than prior works. Even a simple model using this type of auxiliary data outperforms state-of-the-art models that rely on standard features of word embedding encodings of class names. These results highlight the usefulness and importance of textual descriptions for ZSL, as well as the relative importance of auxiliary data type compared to algorithmic progress. Our experimental results also show that standard zero-shot learning approaches generalize poorly across categories of classes.

updated: Wed Mar 17 2021 14:06:56 GMT+0000 (UTC)

published: Wed Mar 17 2021 14:06:56 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト