CentripetalText: An Efficient Text Instance Representation for Scene Text Detection

Tao Sheng; Jie Chen; Zhouhui Lian

CentripetalText：シーンテキスト検出のための効率的なテキストインスタンス表現

シーンテキストの検出は、テキストの曲率、向き、アスペクト比が異なるため、依然として大きな課題です。このタスクで最も難しい問題の1つは、任意の形状のテキストインスタンスをどのように表現するかです。不規則なテキストを柔軟にモデル化するために多くの方法が提案されていますが、それらのほとんどは単純さと堅牢性を失っています。それらの複雑な後処理とディラックのデルタ分布の下での回帰は、検出性能と一般化能力を損ないます。この論文では、CentripetalText（CT）という名前の効率的なテキストインスタンス表現を提案します。これは、テキストインスタンスをテキストカーネルと求心シフトの組み合わせに分解します。具体的には、求心シフトを利用してピクセル集約を実装し、外部テキストピクセルを内部テキストカーネルに導きます。緩和操作は求心シフトの密な回帰に統合され、特定の値ではなく範囲で正しい予測を可能にします。私たちの方法でのテキスト輪郭の便利な再構成と予測誤差の許容度は、それぞれ高い検出精度と速い推論速度を保証します。さらに、テキスト検出器をプロポーザル生成モジュール、つまりCentripetalTextプロポーザルネットワークに縮小し、Mask TextSpotter v3のセグメンテーションプロポーザルネットワークを置き換えて、より正確なプロポーザルを作成します。私たちの方法の有効性を検証するために、湾曲したテキストデータセットと多方向のテキストデータセットの両方を含む、いくつかの一般的に使用されるシーンテキストベンチマークで実験を行います。シーンテキスト検出のタスクでは、他の既存の方法と比較して、優れたパフォーマンスまたは競争力のあるパフォーマンスを実現します。たとえば、Total-Textで40.0 FPSでFメジャーが86.3％、MSRA-TD500で34.8 FPSでFメジャーが86.1％です。、など。エンドツーエンドのシーンテキスト認識のタスクでは、このメソッドは、Total-TextでMask TextSpotter v3を1.1％上回っています。

Scene text detection remains a grand challenge due to the variation in text curvatures, orientations, and aspect ratios. One of the hardest problems in this task is how to represent text instances of arbitrary shapes. Although many methods have been proposed to model irregular texts in a flexible manner, most of them lose simplicity and robustness. Their complicated post-processings and the regression under Dirac delta distribution undermine the detection performance and the generalization ability. In this paper, we propose an efficient text instance representation named CentripetalText (CT), which decomposes text instances into the combination of text kernels and centripetal shifts. Specifically, we utilize the centripetal shifts to implement pixel aggregation, guiding the external text pixels to the internal text kernels. The relaxation operation is integrated into the dense regression for centripetal shifts, allowing the correct prediction in a range instead of a specific value. The convenient reconstruction of text contours and the tolerance of prediction errors in our method guarantee the high detection accuracy and the fast inference speed, respectively. Besides, we shrink our text detector into a proposal generation module, namely CentripetalText Proposal Network, replacing Segmentation Proposal Network in Mask TextSpotter v3 and producing more accurate proposals. To validate the effectiveness of our method, we conduct experiments on several commonly used scene text benchmarks, including both curved and multi-oriented text datasets. For the task of scene text detection, our approach achieves superior or competitive performance compared to other existing methods, e.g., F-measure of 86.3% at 40.0 FPS on Total-Text, F-measure of 86.1% at 34.8 FPS on MSRA-TD500, etc. For the task of end-to-end scene text recognition, our method outperforms Mask TextSpotter v3 by 1.1% on Total-Text.

updated: Tue Oct 26 2021 17:09:52 GMT+0000 (UTC)

published: Tue Jul 13 2021 09:34:18 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト