Complementary Feature Enhanced Network with Vision Transformer for Image Dehazing

Dong Zhao; Jia Li; Hongyu Li; Long Xu

画像の曇り除去のためのビジョントランスフォーマーを備えた補完機能強化ネットワーク

従来のCNNベースのデヘイズモデルには、デヘイズフレームワーク（解釈可能性が制限されている）と畳み込み層（コンテンツに依存せず、長距離の依存関係情報を学習するには効果がない）という2つの重要な問題があります。この論文では、最初に、補完的な機能がいくつかの補完的なサブタスクによって学習され、次に一緒になって主要なタスクのパフォーマンスを向上させる、新しい補完的な機能強化フレームワークを提案します。新しいフレームワークの顕著な利点の1つは、意図的に選択された補完的なタスクが、弱く依存する補完的な機能の学習に集中でき、ネットワークの反復的で効果のない学習を回避できることです。このようなフレームワークに基づいて、新しいヘイズ除去ネットワークを設計します。具体的には、固有の画像分解を補完的なタスクとして選択します。ここでは、反射率とシェーディングの予測サブタスクを使用して、色とテクスチャの補完的な特徴を抽出します。これらの補完的な機能を効果的に集約するために、画像の曇り除去にさらに役立つ機能を選択するための補完的な機能選択モジュール（CFSM）を提案します。さらに、Hybrid Local-Global Vision Transformer（HyLoG-ViT）という名前の新しいバージョンのビジョントランスフォーマーブロックを導入し、それをデヘイズネットワークに組み込みます。 HyLoG-ViTブロックは、ローカルおよびグローバルの依存関係をキャプチャするために使用されるローカルおよびグローバルのビジョントランスフォーマーパスで構成されます。その結果、HyLoG-ViTはネットワークに局所性を導入し、グローバルおよび長期の依存関係をキャプチャします。均質、不均質、および夜間のヘイズ除去タスクに関する広範な実験により、提案されたヘイズ除去ネットワークは、CNNベースのヘイズ除去モデルと同等またはそれ以上のパフォーマンスを達成できることが明らかになりました。

Conventional CNNs-based dehazing models suffer from two essential issues: the dehazing framework (limited in interpretability) and the convolution layers (content-independent and ineffective to learn long-range dependency information). In this paper, firstly, we propose a new complementary feature enhanced framework, in which the complementary features are learned by several complementary subtasks and then together serve to boost the performance of the primary task. One of the prominent advantages of the new framework is that the purposively chosen complementary tasks can focus on learning weakly dependent complementary features, avoiding repetitive and ineffective learning of the networks. We design a new dehazing network based on such a framework. Specifically, we select the intrinsic image decomposition as the complementary tasks, where the reflectance and shading prediction subtasks are used to extract the color-wise and texture-wise complementary features. To effectively aggregate these complementary features, we propose a complementary features selection module (CFSM) to select the more useful features for image dehazing. Furthermore, we introduce a new version of vision transformer block, named Hybrid Local-Global Vision Transformer (HyLoG-ViT), and incorporate it within our dehazing networks. The HyLoG-ViT block consists of the local and the global vision transformer paths used to capture local and global dependencies. As a result, the HyLoG-ViT introduces locality in the networks and captures the global and long-range dependencies. Extensive experiments on homogeneous, non-homogeneous, and nighttime dehazing tasks reveal that the proposed dehazing network can achieve comparable or even better performance than CNNs-based dehazing models.

updated: Wed Jan 05 2022 03:05:45 GMT+0000 (UTC)

published: Wed Sep 15 2021 06:13:22 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト