GKNet: grasp keypoint network for grasp candidates detection

Ruinian Xu; Fu-Jen Chu; Patricio A. Vela

GKNet：把握候補者検出のための把握キーポイントネットワーク

現代の把握検出アプローチは、センサーとオブジェクトモデルの不確実性に対するロバスト性を実現するためにディープラーニングを採用しています。 2つの主要なアプローチは、把握品質スコアリングまたはアンカーベースの把握認識ネットワークのいずれかを設計します。この論文は、画像空間におけるキーポイント検出として扱うことにより、検出を把握するための異なるアプローチを提示します。ディープネットワークは、各把握候補をキーポイントのペアとして検出し、コーナーポイントのトリプレットまたはカルテットではなく、把握表現g = {x、y、w、θ} Tに変換できます。キーポイントをペアにグループ化して検出の難易度を下げると、パフォーマンスが向上します。キーポイント間の依存関係のキャプチャを促進するために、非ローカルモジュールがネットワーク設計に組み込まれています。離散的かつ連続的な方向予測に基づく最終的なフィルタリング戦略は、誤った対応を取り除き、把握検出のパフォーマンスをさらに向上させます。ここで紹介するアプローチであるGKNetは、Cornellと簡略化されたJacquardデータセットの精度と速度のバランスが取れています（41.67および23.26 fpsで96.9％および98.39％）。マニピュレータの追跡実験では、静的把持、動的把持、さまざまなカメラアングルでの把持、ビンピッキングの4種類の妨害実験を使用して、GKNetを評価します。 GKNetは、静的および動的な把握実験で参照ベースラインを上回り、さまざまなカメラの視点と適度な乱雑さに対する堅牢性を示しています。結果は、把握キーポイントが予想される迷惑要因にロバスト性を提供する深い把握ネットワークの効果的な出力表現であるという仮説を確認します。

Contemporary grasp detection approaches employ deep learning to achieve robustness to sensor and object model uncertainty. The two dominant approaches design either grasp-quality scoring or anchor-based grasp recognition networks. This paper presents a different approach to grasp detection by treating it as keypoint detection in image-space. The deep network detects each grasp candidate as a pair of keypoints, convertible to the grasp representationg = {x, y, w, θ} T , rather than a triplet or quartet of corner points. Decreasing the detection difficulty by grouping keypoints into pairs boosts performance. To promote capturing dependencies between keypoints, a non-local module is incorporated into the network design. A final filtering strategy based on discrete and continuous orientation prediction removes false correspondences and further improves grasp detection performance. GKNet, the approach presented here, achieves a good balance between accuracy and speed on the Cornell and the abridged Jacquard datasets (96.9% and 98.39% at 41.67 and 23.26 fps). Follow-up experiments on a manipulator evaluate GKNet using 4 types of grasping experiments reflecting different nuisance sources: static grasping, dynamic grasping, grasping at varied camera angles, and bin picking. GKNet outperforms reference baselines in static and dynamic grasping experiments while showing robustness to varied camera viewpoints and moderate clutter. The results confirm the hypothesis that grasp keypoints are an effective output representation for deep grasp networks that provide robustness to expected nuisance factors.

updated: Wed Dec 15 2021 00:09:26 GMT+0000 (UTC)

published: Wed Jun 16 2021 00:34:55 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト