Characterizing Out-of-Distribution Error via Optimal Transport

Yuzhe Lu; Yilong Qin; Runtian Zhai; Andrew Shen; Ketong Chen; Zhenlin Wang; Soheil Kolouri; Simon Stepputtis; Joseph Campbell; Katia Sycara

最適な輸送による分布外エラーの特徴付け

配布外 (OOD) データは、展開された機械学習モデルに深刻な課題を引き起こすため、ラベルなしで OOD データに対するモデルのパフォーマンスを予測する方法は、機械学習の安全性にとって重要です。これまでの研究で多くの方法が提案されてきましたが、それらは実際の誤差を過小評価することが多く、場合によっては大幅に誤差が生じるため、実際のタスクへの適用性に大きな影響を与えます。この研究では、この過小評価の重要な指標として、擬似ラベルのシフト、つまり予測された OOD ラベル分布と真の OOD ラベル分布の差を特定しました。この観察に基づいて、最適伝送理論である Confidence Optimal Transport (COT) を利用してモデルのパフォーマンスを推定する新しい方法を導入し、それが擬似ラベルシフトの存在下でより堅牢な誤差推定を提供する可能性があることを示します。さらに、経験に基づいた COT である Confidence Optimal Transport with Thresholding (COTT) を導入します。これは、個々の転送コストにしきい値を適用し、COT の誤差推定の精度をさらに向上させます。私たちは、合成、新規亜集団、自然など、さまざまなタイプの分布シフトを誘発するさまざまな標準ベンチマークで COT と COTT を評価し、私たちのアプローチが既存の最先端の手法を最大 3 倍で大幅に上回るパフォーマンスを示すことを示しています。予測誤差が低くなります。

Out-of-distribution (OOD) data poses serious challenges in deployed machine learning models, so methods of predicting a model's performance on OOD data without labels are important for machine learning safety. While a number of methods have been proposed by prior work, they often underestimate the actual error, sometimes by a large margin, which greatly impacts their applicability to real tasks. In this work, we identify pseudo-label shift, or the difference between the predicted and true OOD label distributions, as a key indicator to this underestimation. Based on this observation, we introduce a novel method for estimating model performance by leveraging optimal transport theory, Confidence Optimal Transport (COT), and show that it provably provides more robust error estimates in the presence of pseudo-label shift. Additionally, we introduce an empirically-motivated variant of COT, Confidence Optimal Transport with Thresholding (COTT), which applies thresholding to the individual transport costs and further improves the accuracy of COT's error estimates. We evaluate COT and COTT on a variety of standard benchmarks that induce various types of distribution shift -- synthetic, novel subpopulation, and natural -- and show that our approaches significantly outperform existing state-of-the-art methods with an up to 3x lower prediction error.

updated: Sat May 27 2023 01:08:15 GMT+0000 (UTC)

published: Thu May 25 2023 01:37:13 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト