Automatic Block-wise Pruning with Auxiliary Gating Structures for Deep Convolutional Neural Networks

Zhaofeng Si; Honggang Qi; Xiaoyu Song

深い畳み込みニューラルネットワークのための補助ゲーティング構造による自動ブロックワイズ剪定

畳み込みニューラルネットワークは、深層学習タスクで普及しています。ただし、モバイルデバイスで作業する場合、コストの問題が大きくなります。ネットワークプルーニングは、このような問題を処理するためのモデル圧縮の効果的な方法です。この論文は、剪定時の基準としてバックボーンネットワーク内のブロックに重要度マークを割り当てる補助ゲーティング構造を備えた新しい構造化ネットワーク剪定方法を提示します。次に、ブロック単位のプルーニングは、提案された投票戦略によって実現されます。これは、チャネル単位のようにモデルを小さな粒度でプルーニングする一般的な方法とは異なります。さらに、パフォーマンスを向上させるための知識蒸留を組み込んだ、提案されたアーキテクチャの3段階のトレーニングスケジューリングを開発します。私たちの実験は、私たちの方法が分類タスクのための最先端の圧縮性能を達成できることを示しています。さらに、私たちのアプローチは、事前にトレーニングされたモデルを提供することにより、他の剪定方法と相乗的に統合できるため、93％以上のFLOPが削減された剪定されていないモデルよりも優れたパフォーマンスを実現します。

Convolutional neural networks are prevailing in deep learning tasks. However, they suffer from massive cost issues when working on mobile devices. Network pruning is an effective method of model compression to handle such problems. This paper presents a novel structured network pruning method with auxiliary gating structures which assigns importance marks to blocks in backbone network as a criterion when pruning. Block-wise pruning is then realized by proposed voting strategy, which is different from prevailing methods who prune a model in small granularity like channel-wise. We further develop a three-stage training scheduling for the proposed architecture incorporating knowledge distillation for better performance. Our experiments demonstrate that our method can achieve state-of-the-arts compression performance for the classification tasks. In addition, our approach can integrate synergistically with other pruning methods by providing pretrained models, thus achieving a better performance than the unpruned model with over 93% FLOPs reduced.

updated: Sat May 07 2022 09:03:32 GMT+0000 (UTC)

published: Sat May 07 2022 09:03:32 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト