S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-bit Neural Networks via Guided Distribution Calibration

Zhiqiang Shen; Zechun Liu; Jie Qin; Lei Huang; Kwang-Ting Cheng; Marios Savvides

S2-BNN：ガイド付き分布キャリブレーションによる自己教師ありリアルネットワークと1ビットニューラルネットワーク間のギャップの橋渡し

以前の研究は、主に実数値ネットワークでの自己教師あり学習を対象としており、多くの有望な結果を達成しています。ただし、より困難なバイナリニューラルネットワーク（BNN）では、このタスクはまだコミュニティで十分に検討されていません。このホワイトペーパーでは、このより難しいシナリオに焦点を当てます。つまり、重みとアクティベーションの両方がバイナリであり、人間が注釈を付けたラベルがない学習ネットワークです。バックボーンネットワークには比較的限られた容量と表現能力しか含まれていないため、一般的に使用される対照的な目的は、競争力のある精度のためにBNNを満足させるものではないことがわかります。したがって、パフォーマンスの大幅な低下を引き起こす既存の自己教師あり手法を直接適用する代わりに、最終的な予測分布で実数値から蒸留バイナリネットワークへの新しいガイドあり学習パラダイムを提示して、損失を最小限に抑え、望ましい精度を取得します。私たちの提案する方法は、BNNで5.5〜15％の絶対ゲインによって単純な対照学習ベースラインを高めることができます。さらに、ラベルなしでトレーニングする場合、BNNが実数値モデルと同様の予測分布を回復することは困難であることを明らかにします。したがって、それらをどのように調整するかが、パフォーマンスの低下に対処するための鍵となります。大規模なImageNetおよびダウンストリームデータセットで広範な実験が行われます。私たちの方法は、単純な対照学習ベースラインを大幅に改善し、多くの主流の教師ありBNN方法にさえ匹敵します。コードはhttps://github.com/szq0214/S2-BNNで入手できます。

Previous studies dominantly target at self-supervised learning on real-valued networks and have achieved many promising results. However, on the more challenging binary neural networks (BNNs), this task has not yet been fully explored in the community. In this paper, we focus on this more difficult scenario: learning networks where both weights and activations are binary, meanwhile, without any human annotated labels. We observe that the commonly used contrastive objective is not satisfying on BNNs for competitive accuracy, since the backbone network contains relatively limited capacity and representation ability. Hence instead of directly applying existing self-supervised methods, which cause a severe decline in performance, we present a novel guided learning paradigm from real-valued to distill binary networks on the final prediction distribution, to minimize the loss and obtain desirable accuracy. Our proposed method can boost the simple contrastive learning baseline by an absolute gain of 5.5~15% on BNNs. We further reveal that it is difficult for BNNs to recover the similar predictive distributions as real-valued models when training without labels. Thus, how to calibrate them is key to address the degradation in performance. Extensive experiments are conducted on the large-scale ImageNet and downstream datasets. Our method achieves substantial improvement over the simple contrastive learning baseline, and is even comparable to many mainstream supervised BNN methods. Code is available at https://github.com/szq0214/S2-BNN.

updated: Mon Jun 21 2021 16:10:03 GMT+0000 (UTC)

published: Wed Feb 17 2021 18:59:28 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト