S4ND: Modeling Images and Videos as Multidimensional Signals Using State Spaces

Eric Nguyen; Karan Goel; Albert Gu; Gordon W. Downs; Preey Shah; Tri Dao; Stephen A. Baccus; Christopher Ré

S4ND: 状態空間を使用した多次元信号としての画像とビデオのモデル化

画像やビデオなどの視覚データは、通常、本質的に連続した多次元信号の離散化としてモデル化されます。既存の連続信号モデルは、視覚 (画像など) データの基になる信号を直接モデル化することにより、この事実を利用しようとします。ただし、これらのモデルは、大規模な画像やビデオの分類などの実際の視覚タスクでは、まだ競争力のあるパフォーマンスを達成できていません。ディープステートスペースモデル (SSM) に関する最近の作業ラインに基づいて、SSM の連続信号モデリング機能を画像やビデオを含む多次元データに拡張する新しい多次元 SSM レイヤーである \method を提案します。 S4ND が 1D、2D、および 3D の大規模なビジュアルデータを連続的な多次元信号としてモデル化できることを示し、Conv2D および自己注意レイヤーを既存の最先端モデルの \method\ レイヤーと単純に交換することで強力なパフォーマンスを発揮することを示します。 . ImageNet-1k では、\method\ は、パッチの 1D シーケンスでトレーニングする場合、Vision Transformer ベースラインのパフォーマンスを 1.5% 超え、2D で画像をモデル化する場合は ConvNeXt と一致します。ビデオの場合、S4ND は HMDB-51 のアクティビティ分類で膨張した 3D ConvNeXt を 4% 改善します。 S4ND は、構成によって解像度不変であるグローバルな連続畳み込みカーネルを暗黙的に学習し、複数の解像度にわたる一般化を可能にする誘導バイアスを提供します。エイリアシングを克服するために S4 への単純な帯域制限の変更を開発することにより、S4ND は強力なゼロショット (トレーニング時には見えない) 解像度パフォーマンスを達成し、8 × 8 でトレーニングされ、32 × 32 でテストされた場合、CIFAR-10 でベースラインの Conv2D を 40% 上回っています。画像。プログレッシブサイズ変更でトレーニングすると、S4ND は高解像度モデルの約 1% 以内になり、トレーニングは 22% 速くなります。

Visual data such as images and videos are typically modeled as discretizations of inherently continuous, multidimensional signals. Existing continuous-signal models attempt to exploit this fact by modeling the underlying signals of visual (e.g., image) data directly. However, these models have not yet been able to achieve competitive performance on practical vision tasks such as large-scale image and video classification. Building on a recent line of work on deep state space models (SSMs), we propose \method, a new multidimensional SSM layer that extends the continuous-signal modeling ability of SSMs to multidimensional data including images and videos. We show that S4ND can model large-scale visual data in 1D, 2D, and 3D as continuous multidimensional signals and demonstrates strong performance by simply swapping Conv2D and self-attention layers with \method\ layers in existing state-of-the-art models. On ImageNet-1k, \method\ exceeds the performance of a Vision Transformer baseline by 1.5% when training with a 1D sequence of patches, and matches ConvNeXt when modeling images in 2D. For videos, S4ND improves on an inflated 3D ConvNeXt in activity classification on HMDB-51 by 4%. S4ND implicitly learns global, continuous convolutional kernels that are resolution invariant by construction, providing an inductive bias that enables generalization across multiple resolutions. By developing a simple bandlimiting modification to S4 to overcome aliasing, S4ND achieves strong zero-shot (unseen at training time) resolution performance, outperforming a baseline Conv2D by 40% on CIFAR-10 when trained on 8 ×8 and tested on 32 ×32 images. When trained with progressive resizing, S4ND comes within ∼1% of a high-resolution model while training 22% faster.

updated: Wed Oct 12 2022 20:55:07 GMT+0000 (UTC)

published: Wed Oct 12 2022 20:55:07 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト