B-SMALL: A Bayesian Neural Network approach to Sparse Model-Agnostic Meta-Learning

Anish Madan; Ranjitha Prasad

B-SMALL：スパースモデルにとらわれないメタ学習へのベイズニューラルネットワークアプローチ

モデルがいくつかのトレーニング例を使用して新しいタスクを推測する、メタ学習としても知られる学習から学習へのパラダイムへの関心が高まっています。最近、メタ学習ベースの方法は、数ショットの分類、回帰、強化学習、およびドメイン適応で広く使用されています。モデルにとらわれないメタ学習（MAML）アルゴリズムは、メタトレーニングフェーズでモデルパラメーターの初期化を取得するよく知られたアルゴリズムです。メタテストフェーズでは、この初期化は勾配降下法を使用して新しいタスクに迅速に適応されます。ただし、メタ学習モデルは、トレーニングタスクが不十分であるために過剰適合する傾向があり、その結果、パラメーター化されたモデルが過剰になり、目に見えないタスクの一般化パフォーマンスが低下します。この論文では、ベイズニューラルネットワークベースのMAMLアルゴリズムを提案します。これをB-SMALLアルゴリズムと呼びます。提案されたフレームワークは、正則化としてスパース化近似KL発散を使用するMAMLの損失関数とともにスパース変分損失項を組み込んでいます。分類タスクと回帰タスクを使用してB-MAMLのパフォーマンスを示し、MAMLを使用してスパース化BNNをトレーニングすると、MAMLアプローチと同等のパフォーマンスを実現しながら、モデルのパラメーターフットプリントが実際に向上することを強調します。また、スパース性とメタ学習が有益な分散型センサーネットワークでのアプローチの適用可能性についても説明します。

There is a growing interest in the learning-to-learn paradigm, also known as meta-learning, where models infer on new tasks using a few training examples. Recently, meta-learning based methods have been widely used in few-shot classification, regression, reinforcement learning, and domain adaptation. The model-agnostic meta-learning (MAML) algorithm is a well-known algorithm that obtains model parameter initialization at meta-training phase. In the meta-test phase, this initialization is rapidly adapted to new tasks by using gradient descent. However, meta-learning models are prone to overfitting since there are insufficient training tasks resulting in over-parameterized models with poor generalization performance for unseen tasks. In this paper, we propose a Bayesian neural network based MAML algorithm, which we refer to as the B-SMALL algorithm. The proposed framework incorporates a sparse variational loss term alongside the loss function of MAML, which uses a sparsifying approximated KL divergence as a regularizer. We demonstrate the performance of B-MAML using classification and regression tasks, and highlight that training a sparsifying BNN using MAML indeed improves the parameter footprint of the model while performing at par or even outperforming the MAML approach. We also illustrate applicability of our approach in distributed sensor networks, where sparsity and meta-learning can be beneficial.

updated: Fri Jan 01 2021 09:19:48 GMT+0000 (UTC)

published: Fri Jan 01 2021 09:19:48 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト