Dual Transfer Learning for Event-based End-task Prediction via Pluggable Event to Image Translation

Lin Wang; Yujeong Chae; Kuk-Jin Yoon

プラグ可能なイベントから画像への変換によるイベントベースのエンドタスク予測のためのデュアルトランスファー学習

イベントカメラは、ピクセルごとの強度の変化を認識し、ダイナミックレンジが高くモーションブラーが少ない非同期イベントストリームを出力する新しいセンサーです。エンコーダーデコーダーのようなネットワークに基づいて、イベントのみをエンドタスク学習、たとえばセマンティックセグメンテーションに使用できることが示されています。ただし、イベントはまばらで、ほとんどがエッジ情報を反映しているため、デコーダーだけに依存して元の詳細を復元することは困難です。さらに、ほとんどの方法は、監視のためにピクセル単位の損失のみに頼っています。これは、まばらなイベントからの視覚的な詳細を十分に活用するには不十分であるため、最適なパフォーマンスが低下します。このホワイトペーパーでは、追加の推論コストを追加することなく、エンドタスクのパフォーマンスを効果的に強化するために、Dual Transfer Learning（DTL）という名前のシンプルで柔軟な2ストリームフレームワークを提案します。提案されたアプローチは、イベントからエンドタスク学習（EEL）ブランチ、イベントから画像変換（EIT）ブランチ、および機能レベルのアフィニティ情報とピクセルレベルの知識を同時に探索する転送学習（TL）モジュールの3つの部分で構成されます。 EELブランチを改善するためのEITブランチ。このシンプルでありながら斬新な方法は、イベントからの強力な表現学習につながり、セマンティックセグメンテーションや深度推定などのエンドタスクのパフォーマンスが大幅に向上することで証明されます。

Event cameras are novel sensors that perceive the per-pixel intensity changes and output asynchronous event streams with high dynamic range and less motion blur. It has been shown that events alone can be used for end-task learning, e.g., semantic segmentation, based on encoder-decoder-like networks. However, as events are sparse and mostly reflect edge information, it is difficult to recover original details merely relying on the decoder. Moreover, most methods resort to pixel-wise loss alone for supervision, which might be insufficient to fully exploit the visual details from sparse events, thus leading to less optimal performance. In this paper, we propose a simple yet flexible two-stream framework named Dual Transfer Learning (DTL) to effectively enhance the performance on the end-tasks without adding extra inference cost. The proposed approach consists of three parts: event to end-task learning (EEL) branch, event to image translation (EIT) branch, and transfer learning (TL) module that simultaneously explores the feature-level affinity information and pixel-level knowledge from the EIT branch to improve the EEL branch. This simple yet novel method leads to strong representation learning from events and is evidenced by the significant performance boost on the end-tasks such as semantic segmentation and depth estimation.

updated: Wed Nov 24 2021 09:18:51 GMT+0000 (UTC)

published: Sat Sep 04 2021 06:49:09 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト