Collaborative Unsupervised Visual Representation Learning from Decentralized Data

Weiming Zhuang; Xin Gan; Yonggang Wen; Shuai Zhang; Shuai Yi

分散データからの協調的教師なし視覚表現学習

教師なし表現学習は、インターネット上で利用可能な一元化されたデータを使用して、卓越したパフォーマンスを達成しました。ただし、プライバシー保護に対する意識の高まりにより、複数の関係者（携帯電話やカメラなど）で爆発的に増大する分散型のラベルなし画像データの共有が制限されています。そのため、当然の問題は、これらのデータを活用して、データのプライバシーを保護しながら、ダウンストリームタスクの視覚的表現を学習する方法です。この問題に対処するために、新しい連合教師なし学習フレームワークであるFedUを提案します。このフレームワークでは、各当事者は、オンラインネットワークとターゲットネットワークを使用した対照学習を使用して、ラベルのないデータからモデルを個別にトレーニングします。次に、中央サーバーはトレーニング済みモデルを集約し、集約されたモデルでクライアントのモデルを更新します。各当事者は生データにしかアクセスできないため、データのプライバシーが保護されます。複数の関係者間で分散されたデータは、通常、独立しておらず、同じように分散されており（IID以外）、パフォーマンスが低下します。この課題に取り組むために、2つのシンプルで効果的な方法を提案します。1）サーバー集約用のオンラインネットワークのエンコーダーのみをアップロードし、集約されたエンコーダーで更新する通信プロトコルを設計します。 2）非IIDによって引き起こされる発散に基づいて、予測子を更新する方法を動的に決定するための新しいモジュールを導入します。予測子は、オンラインネットワークの他のコンポーネントです。広範な実験とアブレーションは、FedUの有効性と重要性を示しています。非IIDデータの線形および半教師あり評価では、片方だけのトレーニングを5％以上、他の方法を14％以上上回っています。

Unsupervised representation learning has achieved outstanding performances using centralized data available on the Internet. However, the increasing awareness of privacy protection limits sharing of decentralized unlabeled image data that grows explosively in multiple parties (e.g., mobile phones and cameras). As such, a natural problem is how to leverage these data to learn visual representations for downstream tasks while preserving data privacy. To address this problem, we propose a novel federated unsupervised learning framework, FedU. In this framework, each party trains models from unlabeled data independently using contrastive learning with an online network and a target network. Then, a central server aggregates trained models and updates clients' models with the aggregated model. It preserves data privacy as each party only has access to its raw data. Decentralized data among multiple parties are normally non-independent and identically distributed (non-IID), leading to performance degradation. To tackle this challenge, we propose two simple but effective methods: 1) We design the communication protocol to upload only the encoders of online networks for server aggregation and update them with the aggregated encoder; 2) We introduce a new module to dynamically decide how to update predictors based on the divergence caused by non-IID. The predictor is the other component of the online network. Extensive experiments and ablations demonstrate the effectiveness and significance of FedU. It outperforms training with only one party by over 5% and other methods by over 14% in linear and semi-supervised evaluation on non-IID data.

updated: Sat Aug 14 2021 08:34:11 GMT+0000 (UTC)

published: Sat Aug 14 2021 08:34:11 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト