Discrete Cosine Transform Network for Guided Depth Map Super-Resolution

Zixiang Zhao; Jiangshe Zhang; Shuang Xu; Chunxia Zhang; Junmin Liu

ガイド付き深度マップ超解像のための離散コサイン変換ネットワーク

ガイド付き深度超解像（GDSR）は、マルチモーダル画像処理のホットトピックです。目標は、高解像度（HR）RGB画像を使用して、エッジとオブジェクトの等高線に関する追加情報を提供し、低解像度の深度マップをHRのものにアップサンプリングできるようにすることです。 RGBテクスチャの過剰転送、クロスモーダル特徴抽出の難しさ、および既存の方法におけるモジュールの不明確な動作メカニズムの問題を解決するために、4つのコンポーネントで構成される高度な離散コサイン変換ネットワーク（DCTNet）を提案します。まず、ペアのRGB /深度画像が半結合特徴抽出モジュールに入力されます。共有畳み込みカーネルはクロスモーダル共通機能を抽出し、プライベートカーネルはそれぞれ固有の機能を抽出します。次に、RGB機能がエッジアテンションメカニズムに入力され、アップサンプリングに役立つエッジが強調表示されます。続いて、離散コサイン変換（DCT）モジュールで、DCTを使用して画像ドメインGDSR用に設計された最適化問題を解決します。次に、ソリューションを拡張して、マルチチャネルRGB /深度機能のアップサンプリングを実装します。これにより、DCTNetの合理性が高まり、従来の方法よりも柔軟で効果的です。最終的な深度予測は、再構成モジュールによって出力されます。多数の定性的および定量的実験は、最先端の方法を超えて、正確でHR深度マップを生成できる私たちの方法の有効性を示しています。一方、モジュールの合理性は、アブレーション実験によっても証明されています。

Guided depth super-resolution (GDSR) is a hot topic in multi-modal image processing. The goal is to use high-resolution (HR) RGB images to provide extra information on edges and object contours, so that low-resolution depth maps can be upsampled to HR ones. To solve the issues of RGB texture over-transferred, cross-modal feature extraction difficulty and unclear working mechanism of modules in existing methods, we propose an advanced Discrete Cosine Transform Network (DCTNet), which is composed of four components. Firstly, the paired RGB/depth images are input into the semi-coupled feature extraction module. The shared convolution kernels extract the cross-modal common features, and the private kernels extract their unique features, respectively. Then the RGB features are input into the edge attention mechanism to highlight the edges useful for upsampling. Subsequently, in the Discrete Cosine Transform (DCT) module, where DCT is employed to solve the optimization problem designed for image domain GDSR. The solution is then extended to implement the multi-channel RGB/depth features upsampling, which increases the rationality of DCTNet, and is more flexible and effective than conventional methods. The final depth prediction is output by the reconstruction module. Numerous qualitative and quantitative experiments demonstrate the effectiveness of our method, which can generate accurate and HR depth maps, surpassing state-of-the-art methods. Meanwhile, the rationality of modules is also proved by ablation experiments.

updated: Wed Apr 14 2021 17:01:03 GMT+0000 (UTC)

published: Wed Apr 14 2021 17:01:03 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト