Self-Supervised Learning from Non-Object Centric Images with a Geometric Transformation Sensitive Architecture

Taeho Kim

幾何学的変換に敏感なアーキテクチャを使用した非オブジェクト中心の画像からの自己教師あり学習

ほとんどの不変ベースの自己教師ありメソッドは、幾何学的変換から不変表現を事前トレーニングして学習するために、単一のオブジェクト中心の画像 (ImageNet 画像など) に依存しています。ただし、画像がオブジェクト中心でない場合、トリミングによって画像のセマンティクスが大幅に変更される可能性があります。さらに、モデルは幾何学的に鈍感な特徴を学習するため、位置情報を取得するのに苦労する可能性があります。このため、幾何学的変換、特に 4 回転、ランダムクロップ、マルチクロップに敏感な機能を学習する幾何学的変換センシティブアーキテクチャを提案します。私たちの方法は、教師の特徴マップのプーリングと回転、および回転の予測を介して、これらの変換に敏感なターゲットを使用することにより、生徒が敏感な特徴を学習することを奨励します。さらに、マルチクロップに鈍感なトレーニングは長期的な依存関係をキャプチャできるため、パッチ対応損失を使用して、長期的な依存関係をキャプチャしながらモデルを敏感にトレーニングします。私たちのアプローチは、幾何学的変換に依存しない表現を学習する他の方法と比較して、事前トレーニングデータとして非オブジェクト中心の画像を使用する場合のパフォーマンスの向上を示しています。 6.1 Acc、3.3 mIoU、3.4 AP^b、および 2.7 AP^m の改善により、画像分類、セマンティックセグメンテーション、検出、インスタンスセグメンテーションなどのタスクで DINO[caron2021emerging] ベースラインを上回っています。コードと事前トレーニング済みのモデルは、次の場所で公開されています。

Most invariance-based self-supervised methods rely on single object-centric images (e.g., ImageNet images) for pretraining, learning invariant representations from geometric transformations. However, when images are not object-centric, the semantics of the image can be significantly altered due to cropping. Furthermore, as the model learns geometrically insensitive features, it may struggle to capture location information. For this reason, we propose a Geometric Transformation Sensitive Architecture that learns features sensitive to geometric transformations, specifically four-fold rotation, random crop, and multi-crop. Our method encourages the student to learn sensitive features by using targets that are sensitive to those transforms via pooling and rotating of the teacher feature map and predicting rotation. Additionally, since training insensitively to multi-crop can capture long-term dependencies, we use patch correspondence loss to train the model sensitively while capturing long-term dependencies. Our approach demonstrates improved performance when using non-object-centric images as pretraining data compared to other methods that learn geometric transformation-insensitive representations. We surpass the DINO[caron2021emerging] baseline in tasks including image classification, semantic segmentation, detection, and instance segmentation with improvements of 6.1 Acc, 3.3 mIoU, 3.4 AP^b, and 2.7 AP^m. Code and pretrained models are publicly available at:

updated: Thu Apr 27 2023 04:05:10 GMT+0000 (UTC)

published: Mon Apr 17 2023 06:32:37 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト