Future Frame Prediction for Robot-assisted Surgery

Xiaojie Gao; Yueming Jin; Zixu Zhao; Qi Dou; Pheng-Ann Heng

ロボット支援手術の将来のフレーム予測

ロボット手術ビデオの将来のフレームを予測することは、手術タスクが複雑なダイナミクスを持っている可能性があることを考えると、興味深い、重要であるが非常に困難な問題です。自然ビデオの将来の予測に関する既存のアプローチは、深層リカレントニューラルネットワーク、オプティカルフロー、潜在空間モデリングなど、決定論的モデルまたは確率論的モデルのいずれかに基づいていました。ただし、外科的シナリオでデュアルアームを備えたロボットの意味のある動きを予測する可能性はこれまで活用されていません。これは通常、自然なシナリオで片腕ロボットの独立した動きを予測するよりも困難です。この論文では、ロボット手術ビデオシーケンスにおける将来のフレーム予測のための三元事前誘導変分オートエンコーダ（TPG-VAE）モデルを提案します。コンテンツ配信に加えて、私たちのモデルは、手術器具の小さな動きを処理するための斬新なモーション配信を学習します。さらに、ジェスチャクラスからの不変の事前情報を生成プロセスに追加して、モデルの潜在空間を制約します。私たちの知る限り、デュアルアームロボットの将来のフレームが一般的なロボットビデオと比較した独自の特性を考慮して予測されるのはこれが初めてです。実験は、私たちのモデルが公開JIGSAWSデータセットの縫合タスクでより安定した現実的な将来のフレーム予測シーンを獲得することを示しています。

Predicting future frames for robotic surgical video is an interesting, important yet extremely challenging problem, given that the operative tasks may have complex dynamics. Existing approaches on future prediction of natural videos were based on either deterministic models or stochastic models, including deep recurrent neural networks, optical flow, and latent space modeling. However, the potential in predicting meaningful movements of robots with dual arms in surgical scenarios has not been tapped so far, which is typically more challenging than forecasting independent motions of one arm robots in natural scenarios. In this paper, we propose a ternary prior guided variational autoencoder (TPG-VAE) model for future frame prediction in robotic surgical video sequences. Besides content distribution, our model learns motion distribution, which is novel to handle the small movements of surgical tools. Furthermore, we add the invariant prior information from the gesture class into the generation process to constrain the latent space of our model. To our best knowledge, this is the first time that the future frames of dual arm robots are predicted considering their unique characteristics relative to general robotic videos. Experiments demonstrate that our model gains more stable and realistic future frame prediction scenes with the suturing task on the public JIGSAWS dataset.

updated: Thu Mar 18 2021 15:12:06 GMT+0000 (UTC)

published: Thu Mar 18 2021 15:12:06 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト