Scalable Penalized Regression for Noise Detection in Learning with Noisy Labels

Yikai Wang; Xinwei Sun; Yanwei Fu

ノイズの多いラベルを使用した学習におけるノイズ検出のためのスケーラブルなペナルティ付き回帰

ノイズの多いトレーニングセットは、通常、ニューラルネットワークの一般化と堅牢性の低下につながります。この論文では、理論的に保証されたノイズの多いラベル検出フレームワークを使用して、ノイズの多いラベルを使用した学習（LNL）のノイズの多いデータを検出および削除することを提案します。具体的には、ペナルティ付き回帰を設計して、ネットワーク機能とワンホットラベル間の線形関係をモデル化します。ここで、ノイズの多いデータは、回帰モデルで解決されたゼロ以外の平均シフトパラメーターによって識別されます。フレームワークを多数のカテゴリとトレーニングデータを含むデータセットにスケーラブルにするために、トレーニングセット全体を小さな断片に分割する分割アルゴリズムを提案します。これは、ペナルティ付き回帰によって並行して解決でき、スケーラブルなペナルティ付き回帰につながります（ SPR）フレームワーク。ノイズの多いデータを正しく識別するために、SPRの非漸近確率条件を提供します。 SPRは、標準の教師ありトレーニングパイプラインのサンプル選択モジュールと見なすことができますが、半教師ありアルゴリズムとさらに組み合わせて、ラベルのないデータとしてのノイズの多いデータのサポートをさらに活用します。いくつかのベンチマークデータセットと実際のノイズの多いデータセットでの実験結果は、フレームワークの有効性を示しています。コードと事前トレーニング済みモデルは、https：//github.com/Yikai-Wang/SPR-LNLでリリースされています。

Noisy training set usually leads to the degradation of generalization and robustness of neural networks. In this paper, we propose using a theoretically guaranteed noisy label detection framework to detect and remove noisy data for Learning with Noisy Labels (LNL). Specifically, we design a penalized regression to model the linear relation between network features and one-hot labels, where the noisy data are identified by the non-zero mean shift parameters solved in the regression model. To make the framework scalable to datasets that contain a large number of categories and training data, we propose a split algorithm to divide the whole training set into small pieces that can be solved by the penalized regression in parallel, leading to the Scalable Penalized Regression (SPR) framework. We provide the non-asymptotic probabilistic condition for SPR to correctly identify the noisy data. While SPR can be regarded as a sample selection module for standard supervised training pipeline, we further combine it with semi-supervised algorithm to further exploit the support of noisy data as unlabeled data. Experimental results on several benchmark datasets and real-world noisy datasets show the effectiveness of our framework. Our code and pretrained models are released at https://github.com/Yikai-Wang/SPR-LNL.

updated: Tue Mar 15 2022 11:09:58 GMT+0000 (UTC)

published: Tue Mar 15 2022 11:09:58 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト