Recur, Attend or Convolve? Frame Dependency Modeling Matters for Cross-Domain Robustness in Action Recognition

Sofia Broomé; Ernest Pokropek; Boyu Li; Hedvig Kjellström

繰り返しますか、出席しますか、それとも畳み込みますか？アクション認識におけるクロスドメインロバスト性のためのフレーム依存性モデリングの問題

今日のほとんどのアクション認識モデルは高度にパラメーター化されており、主に空間的に異なるクラスを持つデータセットで評価されます。単一画像の以前の結果は、2D畳み込みニューラルネットワーク（CNN）が、さまざまなコンピュータービジョンタスクの形状ではなくテクスチャに偏る傾向があり（Geirhos et al。、2019）、一般化を減らすことを示しています。まとめると、これは、大規模なビデオモデルが、時間の経過とともに関連する形状を追跡し、それらの動きから一般化可能なセマンティクスを推測するのではなく、疑似相関を学習するという疑いを引き起こします。時間の経過とともに視覚パターンを学習するときにパラメーターの急増を回避する自然な方法は、時間軸全体の繰り返しを利用することです。この記事では、反復、注意ベース、畳み込みビデオモデルのクロスドメインロバストネスをそれぞれ経験的に研究し、このロバストネスがフレーム依存性モデリングの影響を受けるかどうかを調査します。私たちの新しい時間的形状データセットは、単一のフレームからは明らかにされない時間的形状全体を一般化する能力を評価するための軽量データセットとして提案されています。パフォーマンスとレイヤー構造を制御する場合、反復モデルは、畳み込みベースおよび注意ベースのモデルよりも、TemporalShapeデータセットで優れたドメイン外一般化能力を示すことがわかります。さらに、私たちの実験は、畳み込みベースおよび注意ベースのモデルが、再発モデルよりもDiving48でより多くのテクスチャバイアスを示すことを示しています。

Most action recognition models today are highly parameterized, and evaluated on datasets with predominantly spatially distinct classes. Previous results for single images have shown that 2D Convolutional Neural Networks (CNNs) tend to be biased toward texture rather than shape for various computer vision tasks (Geirhos et al., 2019), reducing generalization. Taken together, this raises suspicion that large video models learn spurious correlations rather than to track relevant shapes over time and infer generalizable semantics from their movement. A natural way to avoid parameter explosion when learning visual patterns over time is to make use of recurrence across the time-axis. In this article, we empirically study the cross-domain robustness for recurrent, attention-based and convolutional video models, respectively, to investigate whether this robustness is influenced by the frame dependency modeling. Our novel Temporal Shape dataset is proposed as a light-weight dataset to assess the ability to generalize across temporal shapes which are not revealed from single frames. We find that when controlling for performance and layer structure, recurrent models show better out-of-domain generalization ability on the Temporal Shape dataset than convolution- and attention-based models. Moreover, our experiments indicate that convolution- and attention-based models exhibit more texture bias on Diving48 than recurrent models.

updated: Wed Dec 22 2021 19:11:53 GMT+0000 (UTC)

published: Wed Dec 22 2021 19:11:53 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト