Deformation Robust Roto-Scale-Translation Equivariant CNNs

Liyao Gao; Guang Lin; Wei Zhu

変形ロバストロトスケール-変換同変CNN

群の対称性を学習プロセスに直接組み込むことは、モデル設計の効果的なガイドラインであることが証明されています。入力に対するグループアクションに共変変換することが保証されている機能を生成することにより、グループ同変畳み込みニューラルネットワーク（G-CNN）は、固有の対称性を持つ学習タスクで大幅に改善された一般化パフォーマンスを実現します。 G-CNNの一般的な理論と実際の実装は、回転変換またはスケーリング変換のいずれかの下で、ただし個別にのみ、平面画像について研究されてきました。この論文では、結合されたグループ畳み込みを介してこれら3つのグループ間で共同で同変を達成することが保証されているロトスケール変換同変CNN（RST-CNN）を提示します。さらに、実際の対称変換が完全になることはめったになく、通常は入力変形の影響を受けるため、入力歪みに対する表現の同変の安定性分析を提供します。これにより、（事前に固定された）低周波数空間での畳み込みフィルターの切り捨てられた拡張が促進されます。モード。結果として得られるモデルは、変形に強いRST同変を確実に達成します。つまり、変換が迷惑データの変形によって「汚染」された場合でも、RST対称性は「ほぼ」保持されます。これは、分布外の一般化にとって特に重要なプロパティです。 MNIST、Fashion-MNIST、およびSTL-10の数値実験は、提案されたモデルが、特に回転とスケーリングの両方の変動がデータ内に存在する小さなデータレジームにおいて、従来技術に比べて顕著な利益をもたらすことを示しています。

Incorporating group symmetry directly into the learning process has proved to be an effective guideline for model design. By producing features that are guaranteed to transform covariantly to the group actions on the inputs, group-equivariant convolutional neural networks (G-CNNs) achieve significantly improved generalization performance in learning tasks with intrinsic symmetry. General theory and practical implementation of G-CNNs have been studied for planar images under either rotation or scaling transformation, but only individually. We present, in this paper, a roto-scale-translation equivariant CNN (RST-CNN), that is guaranteed to achieve equivariance jointly over these three groups via coupled group convolutions. Moreover, as symmetry transformations in reality are rarely perfect and typically subject to input deformation, we provide a stability analysis of the equivariance of representation to input distortion, which motivates the truncated expansion of the convolutional filters under (pre-fixed) low-frequency spatial modes. The resulting model provably achieves deformation-robust RST equivariance, i.e., the RST symmetry is still "approximately" preserved when the transformation is "contaminated" by a nuisance data deformation, a property that is especially important for out-of-distribution generalization. Numerical experiments on MNIST, Fashion-MNIST, and STL-10 demonstrate that the proposed model yields remarkable gains over prior arts, especially in the small data regime where both rotation and scaling variations are present within the data.

updated: Mon Nov 22 2021 03:58:24 GMT+0000 (UTC)

published: Mon Nov 22 2021 03:58:24 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト