GKNet: grasp keypoint network for grasp candidates detection

Ruinian Xu; Fu-Jen Chu; Patricio A. Vela

GKNet：把握候補者検出のための把握キーポイントネットワーク

現代の把握検出アプローチでは、ディープラーニングを使用して、センサーとオブジェクトモデルの不確実性に対する堅牢性を実現しています。 2つの主要なアプローチは、把握品質スコアリングまたはアンカーベースの把握認識ネットワークのいずれかを設計します。このホワイトペーパーでは、検出をキーポイント検出として扱うことにより、検出を把握するための別のアプローチを紹介します。ディープネットワークは、各把握候補をキーポイントのペアとして検出し、コーナーポイントのトリプレットまたはカルテットではなく、把握表現g = {x、y、w、θ} ^ Tに変換できます。キーポイントをペアにグループ化して検出の難易度を下げると、パフォーマンスが向上します。キーポイント間の依存関係をさらに促進するために、一般的な非ローカルモジュールが提案された学習フレームワークに組み込まれています。離散的かつ連続的な方向予測に基づく最終的なフィルタリング戦略は、誤った対応を取り除き、把握検出のパフォーマンスをさらに向上させます。ここで紹介するアプローチであるGKNetは、Cornellと簡略化されたJacquardデータセットで精度と速度の最適なバランスを実現します（41.67および23.26 fpsで96.9％および98.39％）。マニピュレータの追跡実験では、静的把持、動的把持、さまざまなカメラアングルでの把持、ビンピッキングの4種類の妨害実験を使用して、GKNetを評価します。 GKNetは、静的および動的な把握実験で参照ベースラインを上回り、さまざまなカメラの視点やビンピッキング実験に対する堅牢性を示しています。結果は、把握キーポイントが予想される迷惑要因にロバスト性を提供する深い把握ネットワークの効果的な出力表現であるという仮説を確認します。

Contemporary grasp detection approaches employ deep learning to achieve robustness to sensor and object model uncertainty. The two dominant approaches design either grasp-quality scoring or anchor-based grasp recognition networks. This paper presents a different approach to grasp detection by treating it as keypoint detection. The deep network detects each grasp candidate as a pair of keypoints, convertible to the grasp representation g = {x, y, w, θ}^T, rather than a triplet or quartet of corner points. Decreasing the detection difficulty by grouping keypoints into pairs boosts performance. To further promote dependencies between keypoints, the general non-local module is incorporated into the proposed learning framework. A final filtering strategy based on discrete and continuous orientation prediction removes false correspondences and further improves grasp detection performance. GKNet, the approach presented here, achieves the best balance of accuracy and speed on the Cornell and the abridged Jacquard dataset (96.9% and 98.39% at 41.67 and 23.26 fps). Follow-up experiments on a manipulator evaluate GKNet using 4 types of grasping experiments reflecting different nuisance sources: static grasping, dynamic grasping, grasping at varied camera angles, and bin picking. GKNet outperforms reference baselines in static and dynamic grasping experiments while showing robustness to varied camera viewpoints and bin picking experiments. The results confirm the hypothesis that grasp keypoints are an effective output representation for deep grasp networks that provide robustness to expected nuisance factors.

updated: Wed Jun 16 2021 00:34:55 GMT+0000 (UTC)

published: Wed Jun 16 2021 00:34:55 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト