Unifying Synergies between Self-supervised Learning and Dynamic Computation

Tarun Krishna; Ayush K Rai; Alexandru Drimbarean; Alan F Smeaton; Kevin McGuinness; Noel E O'Connor

自己教師あり学習と動的計算の相乗効果の統合

自己教師あり学習 (SSL) アプローチは、いくつかのコンピュータービジョンベンチマークで教師あり学習のパフォーマンスをエミュレートすることで、大きな進歩を遂げました。ただし、これにはモデルサイズが大幅に大きくなり、計算コストの高いトレーニング戦略が必要になり、最終的には推論時間が長くなり、リソースに制約のある産業環境では実用的ではなくなります。知識蒸留 (KD)、動的計算 (DC)、プルーニングなどの手法は、軽量のサブネットワークを取得するためによく使用されます。これには通常、大規模な事前トレーニング済みモデルの微調整の複数のエポックが含まれるため、計算がより困難になります。この作業では、SSL と DC パラダイム間の相互作用に関する新しい視点を提案します。これを活用して、高密度およびゲート付き (スパース/軽量) サブネットワークをゼロから同時に学習し、精度と効率のトレードオフを提供し、したがって、アプリケーション固有の産業環境向けの汎用的で多目的なアーキテクチャです。私たちの研究は全体的に建設的なメッセージを伝えています.CIFAR-10、STL-10、CIFAR-100、およびImageNet-100のいくつかの画像分類ベンチマークに関する徹底的な実験は、提案されたトレーニング戦略が高密度で対応する疎なサブネットワークを提供することを示しています。バニラの自己監視設定と比較して同等の（同等の）パフォーマンスですが、さまざまなターゲット予算の下で FLOP に関して計算が大幅に削減されます。

Self-supervised learning (SSL) approaches have made major strides forward by emulating the performance of their supervised counterparts on several computer vision benchmarks. This, however, comes at a cost of substantially larger model sizes, and computationally expensive training strategies, which eventually lead to larger inference times making it impractical for resource constrained industrial settings. Techniques like knowledge distillation (KD), dynamic computation (DC), and pruning are often used to obtain a lightweight sub-network, which usually involves multiple epochs of fine-tuning of a large pre-trained model, making it more computationally challenging. In this work we propose a novel perspective on the interplay between SSL and DC paradigms that can be leveraged to simultaneously learn a dense and gated (sparse/lightweight) sub-network from scratch offering a good accuracy-efficiency trade-off, and therefore yielding a generic and multi-purpose architecture for application specific industrial settings. Our study overall conveys a constructive message: exhaustive experiments on several image classification benchmarks: CIFAR-10, STL-10, CIFAR-100, and ImageNet-100, demonstrates that the proposed training strategy provides a dense and corresponding sparse sub-network that achieves comparable (on-par) performance compared with the vanilla self-supervised setting, but at a significant reduction in computation in terms of FLOPs under a range of target budgets.

updated: Sun Jan 22 2023 17:12:58 GMT+0000 (UTC)

published: Sun Jan 22 2023 17:12:58 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト