GateHUB: Gated History Unit with Background Suppression for Online Action Detection

Junwen Chen; Gaurav Mittal; Ye Yu; Yu Kong; Mei Chen

GateHUB：オンラインアクション検出のためのバックグラウンド抑制を備えたゲート付き履歴ユニット

オンラインアクション検出は、ストリーミングビデオで発生するとすぐにアクションを予測するタスクです。主な課題は、モデルが未来にアクセスできず、予測を行うために履歴、つまりこれまでに観察されたフレームのみに依存する必要があることです。したがって、現在のフレームの予測により有益な履歴の部分を強調することが重要です。 GateHUB、背景抑制を備えたゲート付き履歴ユニットを紹介します。これは、現在のフレーム予測にどれほど有益であるかに従って、履歴の一部を強化または抑制するための新しい位置誘導ゲート交差注意メカニズムを備えています。 GateHUBはさらに、将来観測されるフレームが利用可能な場合にそれを使用することにより、履歴機能をより有益なものにするために、Future-augmented History（FaH）を提案します。単一の統合フレームワークで、GateHUBは、トランスフォーマーの長距離時間モデリングの機能と、関連情報を選択的にエンコードする反復モデルの機能を統合します。 GateHUBは、アクションフレームによく似た誤検知のバックグラウンドフレームをさらに軽減するためのバックグラウンド抑制目標も導入しています。 3つのベンチマークデータセット、THUMOS、TVSeries、およびHDDの広範な検証は、GateHUBが既存のすべての方法を大幅に上回り、既存の最良の作業よりも効率的であることを示しています。さらに、フローフリーバージョンのGateHUBは、予測にRGBとオプティカルフローの両方の情報を必要とする既存のすべての方法と比較して、2.8倍高いフレームレートでより高いまたは近い精度を達成できます。

Online action detection is the task of predicting the action as soon as it happens in a streaming video. A major challenge is that the model does not have access to the future and has to solely rely on the history, i.e., the frames observed so far, to make predictions. It is therefore important to accentuate parts of the history that are more informative to the prediction of the current frame. We present GateHUB, Gated History Unit with Background Suppression, that comprises a novel position-guided gated cross-attention mechanism to enhance or suppress parts of the history as per how informative they are for current frame prediction. GateHUB further proposes Future-augmented History (FaH) to make history features more informative by using subsequently observed frames when available. In a single unified framework, GateHUB integrates the transformer's ability of long-range temporal modeling and the recurrent model's capacity to selectively encode relevant information. GateHUB also introduces a background suppression objective to further mitigate false positive background frames that closely resemble the action frames. Extensive validation on three benchmark datasets, THUMOS, TVSeries, and HDD, demonstrates that GateHUB significantly outperforms all existing methods and is also more efficient than the existing best work. Furthermore, a flow-free version of GateHUB is able to achieve higher or close accuracy at 2.8x higher frame rate compared to all existing methods that require both RGB and optical flow information for prediction.

updated: Thu Jun 09 2022 17:59:44 GMT+0000 (UTC)

published: Thu Jun 09 2022 17:59:44 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト