A Unified Batch Selection Policy for Active Metric Learning

Priyadarshini K; Siddhartha Chaudhuri; Vivek Borkar; Subhasis Chaudhuri

アクティブメトリック学習のための統一されたバッチ選択ポリシー

アクティブメトリック学習は、ある入力ドメインでメトリックの学習モデルを可能な限り迅速に改善するために、注釈を付けるためにトレーニングデータの有用性の高いバッチ（通常は順序付けられたトリプレット）を段階的に選択する問題です。バッチ内の各トリプレットの有益性を独立して評価する標準的なアプローチは、多くの冗長なトリプレットを含む相関性の高いバッチの影響を受けやすく、したがって全体的な有用性が低くなります。最近の研究kumari2020batchは、メトリック学習のためのバッチ非相関戦略を提案していますが、一度に2つのトリプレット間の相関を推定するためにアドホックヒューリスティックに依存しています。最大エントロピー原理を活用して、事前の制約の特定のセットに対する三重項分布の最小バイアス推定値を学習する、新しいバッチアクティブメトリック学習方法を提示します。トリプレット間の冗長性を回避するために、私たちの方法では、結合エントロピーが最大のバッチをまとめて選択します。これにより、情報量と多様性の両方が同時に取得されます。結合エントロピー関数の劣モジュラ性を利用して、グラム・シュミット直交化に基づく効率的な欲張りアルゴリズムを使用して扱いやすい解を構築します。これは、おそらく（1-1e）最適です。私たちのアプローチは、トリプレットのバッチ全体の有益性と多様性のバランスをとる統一スコアを定義する最初のバッチアクティブメトリック学習方法です。いくつかの実際のデータセットを使った実験は、私たちのアルゴリズムが堅牢で、さまざまなアプリケーションや入力モダリティにうまく一般化され、常に最先端を上回っていることを示しています。

Active metric learning is the problem of incrementally selecting high-utility batches of training data (typically, ordered triplets) to annotate, in order to progressively improve a learned model of a metric over some input domain as rapidly as possible. Standard approaches, which independently assess the informativeness of each triplet in a batch, are susceptible to highly correlated batches with many redundant triplets and hence low overall utility. While a recent work kumari2020batch proposes batch-decorrelation strategies for metric learning, they rely on ad hoc heuristics to estimate the correlation between two triplets at a time. We present a novel batch active metric learning method that leverages the Maximum Entropy Principle to learn the least biased estimate of triplet distribution for a given set of prior constraints. To avoid redundancy between triplets, our method collectively selects batches with maximum joint entropy, which simultaneously captures both informativeness and diversity. We take advantage of the submodularity of the joint entropy function to construct a tractable solution using an efficient greedy algorithm based on Gram-Schmidt orthogonalization that is provably ( 1 - 1e )-optimal. Our approach is the first batch active metric learning method to define a unified score that balances informativeness and diversity for an entire batch of triplets. Experiments with several real-world datasets demonstrate that our algorithm is robust, generalizes well to different applications and input modalities, and consistently outperforms the state-of-the-art.

updated: Sun Aug 01 2021 07:34:14 GMT+0000 (UTC)

published: Mon Feb 15 2021 06:55:17 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト