Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal Attention

Katsuyuki Nakamura; Hiroki Ohashi; Mitsuhiro Okada

動的モーダル注意を伴うセンサー拡張自己中心性ビデオキャプション

ビデオまたはビデオキャプションを自動的に記述することは、マルチメディア分野で広く研究されてきました。この論文は、センサー増強エゴセントリックビデオキャプションの新しいタスク、MMACキャプションと呼ばれるそのための新しく構築されたデータセット、およびビデオおよびモーションセンサーのマルチモーダルデータ、または慣性測定ユニットを効果的に利用する新しく提案されたタスクの方法を提案します。（IMU）。従来のビデオキャプションタスクは、固定カメラの視野が限られているため、人間の活動の詳細な説明を処理するのが困難ですが、エゴセントリックビジョンは、はるかに近いものに基づいて人間の活動のより詳細な説明を生成するために使用される可能性が高くなります見る。さらに、ウェアラブルセンサーのデータを補助情報として利用して、エゴセントリックビジョンに固有の問題（モーションブラー、自己閉塞、カメラ範囲外のアクティビティ）を軽減します。コンテキスト情報を考慮し、より注意が必要なモダリティを動的に決定する注意メカニズムに基づいて、センサーデータをビデオデータと組み合わせて有効に活用する方法を提案します。提案されたセンサー融合方法をMMACキャプションデータセットの強力なベースラインと比較し、センサーデータを自己中心性ビデオデータの補足情報として使用することが有益であり、提案された方法が強力なベースラインを上回り、提案された有効性を示していることを発見しました。方法。

Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it called MMAC Captions, and a method for the newly proposed task that effectively utilizes multi-modal data of video and motion sensors, or inertial measurement units (IMUs). While conventional video captioning tasks have difficulty in dealing with detailed descriptions of human activities due to the limited view of a fixed camera, egocentric vision has greater potential to be used for generating the finer-grained descriptions of human activities on the basis of a much closer view. In addition, we utilize wearable-sensor data as auxiliary information to mitigate the inherent problems in egocentric vision: motion blur, self-occlusion, and out-of-camera-range activities. We propose a method for effectively utilizing the sensor data in combination with the video data on the basis of an attention mechanism that dynamically determines the modality that requires more attention, taking the contextual information into account. We compared the proposed sensor-fusion method with strong baselines on the MMAC Captions dataset and found that using sensor data as supplementary information to the egocentric-video data was beneficial, and that our proposed method outperformed the strong baselines, demonstrating the effectiveness of the proposed method.

updated: Tue Sep 07 2021 09:22:09 GMT+0000 (UTC)

published: Tue Sep 07 2021 09:22:09 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト