WinDB: HMD-free and Distortion-free Panoptic Video Fixation Learning

Guotao Wang; Chenglizhao Chen; Aimin Hao; Hong Qin; Deng-Ping Fan

WinDB: HMD フリーで歪みのないパノプティックビデオ固定学習

現在まで、パノプティックビデオで注視収集を実行するために広く採用されている方法は、ヘッドマウントディスプレイ (HMD) に基づいており、HMD を装着して特定のパノプティックシーンを自由に探索しながら、参加者の注視が収集されます。ただし、この広く使用されているデータ収集方法は、断続的な顕著なイベントが含まれている場合に、特定のパノプティック内のどの領域が最も重要であるかを正確に予測するディープモデルをトレーニングするには不十分です。その主な理由は、参加者がパノプティックシーン全体を常に探索するために頭を回転させ続けることができないため、HMD を使用して注視点を収集するときに常に「ブラインドズーム」が存在することです。その結果、収集された注視点は一部の局所的なビューに閉じ込められる傾向があり、残りの領域は「ブラインドズーム」のままになります。したがって、ローカルビューを蓄積する HMD ベースの方法を使用して収集された注視データは、複雑なパノラマシーンの全体的なグローバルな重要性を正確に表すことができません。このペーパーでは、HMD を必要とせず、ブラインドズームフリーのパノプティックビデオ用の Dynamic Blurring (WinDB) 固定コレクションアプローチを備えた補助ウィンドウを紹介します。したがって、収集された注目度は、地域ごとの重要度をよく反映することができる。 WinDB アプローチを使用して、225 以上のカテゴリをカバーする 300 個のパノプティッククリップを含む新しい PanopticVideo-300 データセットをリリースしました。さらに、PanopticVideo-300 を最大限に活用して、ブラインドズームフリーの属性誘発注視移動問題を処理するためのシンプルなベースライン設計を提示しました。

To date, the widely-adopted way to perform fixation collection in panoptic video is based on a head-mounted display (HMD), where participants' fixations are collected while wearing an HMD to explore the given panoptic scene freely. However, this widely-used data collection method is insufficient for training deep models to accurately predict which regions in a given panoptic are most important when it contains intermittent salient events. The main reason is that there always exist "blind zooms" when using HMD to collect fixations since the participants cannot keep spinning their heads to explore the entire panoptic scene all the time. Consequently, the collected fixations tend to be trapped in some local views, leaving the remaining areas to be the "blind zooms". Therefore, fixation data collected using HMD-based methods that accumulate local views cannot accurately represent the overall global importance of complex panoramic scenes. This paper introduces the auxiliary Window with a Dynamic Blurring (WinDB) fixation collection approach for panoptic video, which doesn't need HMD and is blind-zoom-free. Thus, the collected fixations can well reflect the regional-wise importance degree. Using our WinDB approach, we have released a new PanopticVideo-300 dataset, containing 300 panoptic clips covering over 225 categories. Besides, we have presented a simple baseline design to take full advantage of PanopticVideo-300 to handle the blind-zoom-free attribute-induced fixation shifting problem.

updated: Sun May 28 2023 09:14:14 GMT+0000 (UTC)

published: Tue May 23 2023 10:25:22 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト