Recur, Attend or Convolve? Frame Dependency Modeling Matters for Cross-Domain Robustness in Action Recognition

Sofia Broomé; Ernest Pokropek; Boyu Li; Hedvig Kjellström

繰り返しますか、出席しますか、それとも畳み込みますか？行動認識におけるクロスドメインロバスト性のためのフレーム依存性モデリングの問題

今日のほとんどの行動認識モデルは高度にパラメーター化されており、主に空間的に異なるクラスを持つデータセットで評価されます。単一画像の以前の結果は、2D畳み込みニューラルネットワーク（CNN）が、さまざまなコンピュータービジョンタスクの形状ではなくテクスチャに偏る傾向があり（Geirhos et al。、2019）、一般化を減らすことを示しています。まとめると、これは、大規模なビデオモデルが、時間の経過とともに関連する形状を追跡し、それらの動きから一般化可能なセマンティクスを推測するのではなく、疑似相関を学習するという疑いを引き起こします。時間の経過とともに視覚パターンを学習するときにパラメーターの急増を回避する自然な方法は、時間軸全体の再発を利用することです。この記事では、さまざまなフレーム依存性モデリング（反復、注意ベース、または3D畳み込み）を使用したモデルのクロスドメインロバスト性を経験的に研究します。単一のフレームからは明らかにされない、時間的構造をキャプチャする能力の軽量で体系的な評価を可能にするために、TemporalShapeデータセットを提供します。パフォーマンスとレイヤー構造を制御する場合、畳み込み反復モデルは、3D畳み込みベースおよび注意ベースのモデルよりもTemporalShapeデータセットで優れたドメイン外一般化能力を示すことがわかります。さらに、私たちの実験は、畳み込みベースおよび注意ベースのモデルが、畳み込み反復モデルよりもDiving48でより多くのテクスチャバイアスを示すことを示しています。

Most action recognition models today are highly parameterized, and evaluated on datasets with predominantly spatially distinct classes. Previous results for single images have shown that 2D Convolutional Neural Networks (CNNs) tend to be biased toward texture rather than shape for various computer vision tasks (Geirhos et al., 2019), reducing generalization. Taken together, this raises suspicion that large video models learn spurious correlations rather than to track relevant shapes over time and infer generalizable semantics from their movement. A natural way to avoid parameter explosion when learning visual patterns over time is to make use of recurrence across the time-axis. In this article, we empirically study the cross-domain robustness of models with different frame dependency modeling (recurrent, attention-based or 3D convolutional). In order to enable a light-weight and systematic assessment of the ability to capture temporal structure, not revealed from single frames, we provide the Temporal Shape dataset. We find that when controlling for performance and layer structure, convolutional-recurrent models show better out-of-domain generalization ability on the Temporal Shape dataset than 3D convolution- and attention-based models. Moreover, our experiments indicate that convolution- and attention-based models exhibit more texture bias on Diving48 than convolutional-recurrent models.

updated: Sun Mar 27 2022 15:41:04 GMT+0000 (UTC)

published: Wed Dec 22 2021 19:11:53 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト