Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation

Hang Gao; Xizhou Zhu; Steve Lin; Jifeng Dai

変形可能カーネル：オブジェクト変形のための効果的な受容野の適応

畳み込みネットワークは、オブジェクトの幾何学的な変動を認識していないため、モデルとデータ容量の非効率的な利用につながります。この問題を克服するために、変形モデリングに関する最近の研究では、セマンティック認識が変形の影響を受けにくいように、データを共通の配置に向けて空間的に再構成しようとしています。これは通常、静的な演算子を、画像フィールドで学習した自由形式のサンプリンググリッドで補強し、受容フィールドを適応させるためのデータとタスクに動的に調整することによって行われます。しかし、受容野を適応させることは実際の目標に到達することはできません。ネットワークにとって本当に重要なのは、「効果的な」受容野（ERF）です。したがって、実行中にERFを直接適合させる他のアプローチを設計するのは自然です。この作業では、受容フィールドに手を加えずにERFを直接適合させることにより、オブジェクトの変形を処理するための斬新かつ一般的な畳み込み演算子のファミリーであるDeformable Kernel（DK）として可能なソリューションをインスタンス化します。この方法の中心にあるのは、オブジェクトの変形を回復するために元のカーネル空間をリサンプリングする機能です。このアプローチは、ERFがデータサンプリングの場所とカーネル値によって厳密に決定されるという理論的洞察により正当化されます。リジッドカーネルの一般的なドロップイン置換としてDKを実装し、その結果が理論に準拠する一連の実証研究を実施します。いくつかのタスクと標準ベースモデルにわたって、このアプローチは、実行時に適応する以前の作業と比較して有利です。さらに、さらなる実験により、以前の研究と直交し補完する作用メカニズムが示唆されています。

Convolutional networks are not aware of an object's geometric variations, which leads to inefficient utilization of model and data capacity. To overcome this issue, recent works on deformation modeling seek to spatially reconfigure the data towards a common arrangement such that semantic recognition suffers less from deformation. This is typically done by augmenting static operators with learned free-form sampling grids in the image space, dynamically tuned to the data and task for adapting the receptive field. Yet adapting the receptive field does not quite reach the actual goal -- what really matters to the network is the "effective" receptive field (ERF), which reflects how much each pixel contributes. It is thus natural to design other approaches to adapt the ERF directly during runtime. In this work, we instantiate one possible solution as Deformable Kernels (DKs), a family of novel and generic convolutional operators for handling object deformations by directly adapting the ERF while leaving the receptive field untouched. At the heart of our method is the ability to resample the original kernel space towards recovering the deformation of objects. This approach is justified with theoretical insights that the ERF is strictly determined by data sampling locations and kernel values. We implement DKs as generic drop-in replacements of rigid kernels and conduct a series of empirical studies whose results conform with our theories. Over several tasks and standard base models, our approach compares favorably against prior works that adapt during runtime. In addition, further experiments suggest a working mechanism orthogonal and complementary to previous works.

updated: Wed Feb 12 2020 07:10:24 GMT+0000 (UTC)

published: Mon Oct 07 2019 17:58:10 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト