Visualizing and Understanding Self-Supervised Vision Learning

Fawaz Sammani; Boris Joukovsky; Nikos Deligiannis

自己管理型ビジョン学習の視覚化と理解

自己教師ありビジョン学習はディープラーニングに革命をもたらし、ドメイン内の次の大きな課題となり、大規模なコンピュータービジョンベンチマークでの教師あり手法とのギャップを急速に埋めています。現在のモデルとトレーニングデータが指数関数的に増加するにつれて、これらのモデルの説明と理解が極めて重要になります。視覚課題の自己監視学習の領域で説明可能な人工知能の問題を研究し、自己監視で訓練されたネットワークとその内部の仕組みを理解する方法を提示します。自己監視ビジョン口実タスクの非常に多様性を考えると、同じ画像の2つのビューから学習するパラダイムを理解することに焦点を絞り、主に口実タスクを理解することを目的としています。私たちの仕事は類似性学習の説明に焦点を当てており、他のすべての口実タスクに簡単に拡張できます。 SimCLRとBarlowTwinsの2つの人気のある自己監視型ビジョンモデルを研究します。これらのモデルを視覚化および理解するための合計6つの方法を開発します。摂動ベースの方法（条件付きオクルージョン、コンテキストに依存しない条件付きオクルージョン、ペアワイズオクルージョン）、インタラクションCAM、機能の視覚化、モデルの違いの視覚化、平均化された変換、ピクセルの侵入。最後に、単一の画像を含む教師あり画像分類システムに合わせて調整されたよく知られた評価指標を、2つの画像が含まれる教師あり学習のドメインに変換することにより、これらの説明を評価します。コードは次の場所にあります：https：//github.com/fawazsammani/xai-ssl

Self-Supervised vision learning has revolutionized deep learning, becoming the next big challenge in the domain and rapidly closing the gap with supervised methods on large computer vision benchmarks. With current models and training data exponentially growing, explaining and understanding these models becomes pivotal. We study the problem of explainable artificial intelligence in the domain of self-supervised learning for vision tasks, and present methods to understand networks trained with self-supervision and their inner workings. Given the huge diversity of self-supervised vision pretext tasks, we narrow our focus on understanding paradigms which learn from two views of the same image, and mainly aim to understand the pretext task. Our work focuses on explaining similarity learning, and is easily extendable to all other pretext tasks. We study two popular self-supervised vision models: SimCLR and Barlow Twins. We develop a total of six methods for visualizing and understanding these models: Perturbation-based methods (conditional occlusion, context-agnostic conditional occlusion and pairwise occlusion), Interaction-CAM, Feature Visualization, Model Difference Visualization, Averaged Transforms and Pixel Invaraince. Finally, we evaluate these explanations by translating well-known evaluation metrics tailored towards supervised image classification systems involving a single image, into the domain of self-supervised learning where two images are involved. Code is at: https://github.com/fawazsammani/xai-ssl

updated: Mon Jun 20 2022 13:01:46 GMT+0000 (UTC)

published: Mon Jun 20 2022 13:01:46 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト