Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer

Xiaojie Gao; Yueming Jin; Yonghao Long; Qi Dou; Pheng-Ann Heng

Trans-SVNet：ハイブリッド埋め込み集約トランスフォーマーを介した手術ビデオからの正確な位相認識

リアルタイムの手術段階認識は、現代の手術室における基本的なタスクです。以前の作品は、時空間順に配置されたアーキテクチャに依存してこのタスクに取り組んでいますが、中間の空間機能のサポート上の利点は考慮されていません。この論文では、外科的ワークフロー分析で初めて、正確な外科的位相認識のために空間的および時間的特徴の無視された補完的効果を再考するためのTransformerを紹介します。当社のハイブリッド埋め込み集約Transformerは、時間的埋め込みシーケンスからの空間情報に基づくアクティブなクエリを可能にすることにより、巧妙に設計された空間的および時間的埋め込みを融合します。さらに重要なことに、私たちのフレームワークは、ハイブリッド埋め込みを並行して処理して、高い推論速度を実現します。私たちの方法は、2つの大きな手術ビデオデータセット、つまりCholec80およびM2CAI16チャレンジデータセットで徹底的に検証されており、91fpsの処理速度で最先端のアプローチを上回っています。

Real-time surgical phase recognition is a fundamental task in modern operating rooms. Previous works tackle this task relying on architectures arranged in spatio-temporal order, however, the supportive benefits of intermediate spatial features are not considered. In this paper, we introduce, for the first time in surgical workflow analysis, Transformer to reconsider the ignored complementary effects of spatial and temporal features for accurate surgical phase recognition. Our hybrid embedding aggregation Transformer fuses cleverly designed spatial and temporal embeddings by allowing for active queries based on spatial information from temporal embedding sequences. More importantly, our framework processes the hybrid embeddings in parallel to achieve a high inference speed. Our method is thoroughly validated on two large surgical video datasets, i.e., Cholec80 and M2CAI16 Challenge datasets, and outperforms the state-of-the-art approaches at a processing speed of 91 fps.

updated: Mon Jul 12 2021 12:18:39 GMT+0000 (UTC)

published: Wed Mar 17 2021 15:12:55 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト