Part-Aware Fine-grained Object Categorization using Weakly Supervised Part Detection Network

Yabin Zhang; Kui Jia; Zhixin Wang

弱教師付き部品検出ネットワークを使用した部品認識細粒度オブジェクト分類

粒度の細かいオブジェクト分類は、同じエントリレベルのオブジェクトカテゴリに属する下位カテゴリのオブジェクトを区別することを目的としています。このタスクは、（1）グラウンドトゥルースラベル付きのトレーニング画像を取得するのが困難であり、（2）異なる下位カテゴリ間のバリエーションが微妙であるという事実のために困難です。オブジェクトインスタンスのローカル部分に、さまざまな下位カテゴリの特性化機能が配置されていることが十分に確立されています。実際、多くの詳細な分類データセットで、慎重な部品注釈が利用可能です。ただし、オブジェクトパーツに手動で注釈を付けるには専門知識が必要です。これは、新しいきめ細かい分類タスクに一般化することも困難です。この作業では、細分化された分類を使用するための識別可能なローカルパーツを検出できる、弱監視パーツ検出ネットワーク（PartNet）を提案します。バニラPartNetは、ベースサブネットワークの上に、上位ネットワークレイヤーの2つの並列ストリームを構築します。それぞれ、ローカルの関心領域（RoIs）の分類確率（下位カテゴリ）と検出確率（指定数の識別部分検出器）のスコアを計算します）。画像レベルの予測は、これらの地域レベルの確率の要素ごとの積を集約することにより得られます。さまざまなRoIのセットをPartNetの入力として生成するために、オブジェクトレベルの提案によるブリッジングを行わずに、差別的なローカルパーツの候補の提案を直接ターゲットとする単純な離散化パーツ提案モジュール（DPP）を提案します。ベンチマークCUB-200-2011およびOxford Flower 102データセットでの実験は、差別的な部分の検出ときめの細かい分類の両方に対して提案された方法の有効性を示しています。特に、真理部の注釈が使用できない場合、CUB-200-2011データセットで新しい最先端のパフォーマンスを実現します。

Fine-grained object categorization aims for distinguishing objects of subordinate categories that belong to the same entry-level object category. The task is challenging due to the facts that (1) training images with ground-truth labels are difficult to obtain, and (2) variations among different subordinate categories are subtle. It is well established that characterizing features of different subordinate categories are located on local parts of object instances. In fact, careful part annotations are available in many fine-grained categorization datasets. However, manually annotating object parts requires expertise, which is also difficult to generalize to new fine-grained categorization tasks. In this work, we propose a Weakly Supervised Part Detection Network (PartNet) that is able to detect discriminative local parts for use of fine-grained categorization. A vanilla PartNet builds on top of a base subnetwork two parallel streams of upper network layers, which respectively compute scores of classification probabilities (over subordinate categories) and detection probabilities (over a specified number of discriminative part detectors) for local regions of interest (RoIs). The image-level prediction is obtained by aggregating element-wise products of these region-level probabilities. To generate a diverse set of RoIs as inputs of PartNet, we propose a simple Discretized Part Proposals module (DPP) that directly targets for proposing candidates of discriminative local parts, with no bridging via object-level proposals. Experiments on the benchmark CUB-200-2011 and Oxford Flower 102 datasets show the efficacy of our proposed method for both discriminative part detection and fine-grained categorization. In particular, we achieve the new state-of-the-art performance on CUB-200-2011 dataset when ground-truth part annotations are not available.

updated: Wed Dec 04 2019 14:32:09 GMT+0000 (UTC)

published: Sat Jun 16 2018 07:08:59 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト