ISP4ML: Understanding the Role of Image Signal Processing in Efficient Deep Learning Vision Systems

Patrick Hansen; Alexey Vilkin; Yury Khrustalev; James Imber; David Hanwell; Matthew Mattina; Paul N. Whatmough

ISP4ML：効率的な深層学習ビジョンシステムにおける画像信号処理の役割の理解

畳み込みニューラルネットワーク（CNN）は現在、さまざまなコンピュータービジョン（CV）システムの主要なコンポーネントです。 ISPは従来、人間にとって魅力的な画像を生成するように設計されていましたが、これらのシステムには通常、画像信号プロセッサ（ISP）が含まれています。 CVシステムでは、ISPの役割が明確ではなく、正確な予測のためにISPの役割がまったく必要であるかどうかも不明です。この作業では、CNN分類タスクにおけるISPの有効性を調査し、予測精度と計算コストの間のシステムレベルのトレードオフを概説します。そのために、構成可能なISPとイメージングセンサーのソフトウェアモデルを構築して、さまざまなISP設定と機能を備えたImageNetでCNNをトレーニングします。 ImageNetの結果は、ISPが幅の異なるMobileNetアーキテクチャで4.6〜12.2％精度を向上させることを示しています。 ResNetを使用した結果は、これらの傾向がより深いネットワークに一般化されることを示しています。典型的なISPのさまざまな処理段階のアブレーション研究により、トーンマッパーは、5.8％の平均精度の向上のみを提供することにより、高ダイナミックレンジ（HDR）画像を操作する際の最も重要な段階であることがわかります。全体として、ISPのメモリと計算コストは同じ精度を達成するために大きなCNNを使用するコストに比べて最小限であるため、ISPはシステム効率にメリットをもたらします。

Convolutional neural networks (CNNs) are now predominant components in a variety of computer vision (CV) systems. These systems typically include an image signal processor (ISP), even though the ISP is traditionally designed to produce images that look appealing to humans. In CV systems, it is not clear what the role of the ISP is, or if it is even required at all for accurate prediction. In this work, we investigate the efficacy of the ISP in CNN classification tasks, and outline the system-level trade-offs between prediction accuracy and computational cost. To do so, we build software models of a configurable ISP and an imaging sensor in order to train CNNs on ImageNet with a range of different ISP settings and functionality. Results on ImageNet show that an ISP improves accuracy by 4.6%-12.2% on MobileNet architectures of different widths. Results using ResNets demonstrate that these trends also generalize to deeper networks. An ablation study of the various processing stages in a typical ISP reveals that the tone mapper is the most significant stage when operating on high dynamic range (HDR) images, by providing 5.8% average accuracy improvement alone. Overall, the ISP benefits system efficiency because the memory and computational costs of the ISP is minimal compared to the cost of using a larger CNN to achieve the same accuracy.

updated: Wed Mar 17 2021 15:15:35 GMT+0000 (UTC)

published: Mon Nov 18 2019 21:05:44 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト