Function-Consistent Feature Distillation

Dongyang Liu; Meina Kan; Shiguang Shan; Xilin Chen

機能一貫性のある特徴の蒸留

特徴抽出により、生徒は教師の中間的な特徴を模倣します。ほとんどすべての既存の機能抽出方法は、教師と生徒の機能間の距離メトリックとして L2 距離またはそのわずかな変形を使用します。ただし、L2 距離はすべての次元に対して等方的ですが、異なる次元でのニューラルネットワークの操作は通常異方性です。つまり、2 ノルムが同じでも中間特徴の次元が異なる摂動は、大きさが大きく異なる最終出力の変化につながります。これを考慮して、教師と生徒の特徴間の類似性は、それらの外観 (つまり、L2 距離) だけに基づいて測定されるべきではなく、より重要なこととして、それらの機能の違い、つまりネットワークの後のレイヤーがどのように機能するかによって測定されるべきであると主張します。それらを読み取り、デコードし、処理します。したがって、教師と生徒の機能間の機能的類似性を明示的に最適化する機能一貫性のある機能抽出 (FCFD) を提案します。 FCFD の核となるアイデアは、教師と生徒の特徴を数値的に類似させるだけでなく、より重要なことに、同じネットワークの後半部分に供給されたときに類似の出力を生成することです。 FCFD を使用すると、生徒は教師をより忠実に模倣し、教師からより多くを学びます。画像分類とオブジェクト検出に関する広範な実験により、FCFD が既存の方法よりも優れていることが実証されています。さらに、FCFD を多くの既存の手法と組み合わせて、さらに高い精度を得ることができます。コードは https://github.com/LiuDongyang6/FCFD で入手できます。

Feature distillation makes the student mimic the intermediate features of the teacher. Nearly all existing feature-distillation methods use L2 distance or its slight variants as the distance metric between teacher and student features. However, while L2 distance is isotropic w.r.t. all dimensions, the neural network's operation on different dimensions is usually anisotropic, i.e., perturbations with the same 2-norm but in different dimensions of intermediate features lead to changes in the final output with largely different magnitude. Considering this, we argue that the similarity between teacher and student features should not be measured merely based on their appearance (i.e., L2 distance), but should, more importantly, be measured by their difference in function, namely how later layers of the network will read, decode, and process them. Therefore, we propose Function-Consistent Feature Distillation (FCFD), which explicitly optimizes the functional similarity between teacher and student features. The core idea of FCFD is to make teacher and student features not only numerically similar, but more importantly produce similar outputs when fed to the later part of the same network. With FCFD, the student mimics the teacher more faithfully and learns more from the teacher. Extensive experiments on image classification and object detection demonstrate the superiority of FCFD to existing methods. Furthermore, we can combine FCFD with many existing methods to obtain even higher accuracy. Our codes are available at https://github.com/LiuDongyang6/FCFD.

updated: Mon Apr 24 2023 05:43:29 GMT+0000 (UTC)

published: Mon Apr 24 2023 05:43:29 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト