Contrastive Learning from Spatio-Temporal Mixed Skeleton Sequences for Self-Supervised Skeleton-Based Action Recognition

Zhan Chen; Hong Liu; Tianyu Guo; Zhengyan Chen; Pinhao Song; Hao Tang

自己監視型スケルトンベースの行動認識のための時空間混合スケルトンシーケンスからの対照学習

対照的な学習を伴う自己監視型スケルトンベースの行動認識は、多くの注目を集めています。最近の文献は、データ拡張と対照的なペアの大規模なセットがそのような表現を学習する上で重要であることを示しています。この論文では、通常のデータ拡張から損失への対照ペアの寄与がトレーニングが進むにつれて小さくなるため、通常の拡張に基づいて対照ペアを直接拡張すると、パフォーマンスの点で限られたリターンしかもたらさないことがわかりました。したがって、対照的な学習のために、ハードな対照的なペアを掘り下げます。新規サンプルを合成することによって多くのタスクのパフォーマンスを向上させる混合増強戦略の成功に動機付けられて、SkeleMixCLRを提案します：ハード対照サンプルを提供することによって現在の対照学習アプローチを補完する時空間スケルトン混合増強（SkeleMix）を備えた対照学習フレームワーク。まず、SkeleMixは、スケルトンデータのトポロジ情報を利用して、トリミングされたスケルトンフラグメント（トリミングされたビュー）と残りのスケルトンシーケンス（切り捨てられたビュー）をランダムに組み合わせることにより、2つのスケルトンシーケンスを混合します。次に、時空間マスクプーリングを適用して、これら2つのビューを機能レベルで分離します。第三に、これら2つのビューで対照的なペアを拡張します。 SkeleMixCLRは、トリミングされたビューと切り捨てられたビューを活用して、グラフの畳み込み操作により相互にコンテキスト情報が含まれるため、豊富なハードコントラストペアを提供します。これにより、モデルはアクション認識のためのより優れたモーション表現を学習できます。 NTU-RGB + D、NTU120-RGB + D、およびPKU-MMDデータセットに関する広範な実験は、SkeleMixCLRが最先端のパフォーマンスを達成することを示しています。コードはhttps://github.com/czhaneva/SkeleMixCLRで入手できます。

Self-supervised skeleton-based action recognition with contrastive learning has attracted much attention. Recent literature shows that data augmentation and large sets of contrastive pairs are crucial in learning such representations. In this paper, we found that directly extending contrastive pairs based on normal augmentations brings limited returns in terms of performance, because the contribution of contrastive pairs from the normal data augmentation to the loss get smaller as training progresses. Therefore, we delve into hard contrastive pairs for contrastive learning. Motivated by the success of mixing augmentation strategy which improves the performance of many tasks by synthesizing novel samples, we propose SkeleMixCLR: a contrastive learning framework with a spatio-temporal skeleton mixing augmentation (SkeleMix) to complement current contrastive learning approaches by providing hard contrastive samples. First, SkeleMix utilizes the topological information of skeleton data to mix two skeleton sequences by randomly combing the cropped skeleton fragments (the trimmed view) with the remaining skeleton sequences (the truncated view). Second, a spatio-temporal mask pooling is applied to separate these two views at the feature level. Third, we extend contrastive pairs with these two views. SkeleMixCLR leverages the trimmed and truncated views to provide abundant hard contrastive pairs since they involve some context information from each other due to the graph convolution operations, which allows the model to learn better motion representations for action recognition. Extensive experiments on NTU-RGB+D, NTU120-RGB+D, and PKU-MMD datasets show that SkeleMixCLR achieves state-of-the-art performance. Codes are available at https://github.com/czhaneva/SkeleMixCLR.

updated: Thu Jul 07 2022 03:18:09 GMT+0000 (UTC)

published: Thu Jul 07 2022 03:18:09 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト