Split, embed and merge: An accurate table structure recognizer

Zhenrong Zhang; Jianshu Zhang; Jun Du

分割、埋め込み、マージ：正確なテーブル構造認識機能

テーブル構造認識のタスクは、テーブルの内部構造を認識することです。これは、マシンにテーブルを理解させるための重要なステップです。ただし、Portable Document Format（PDF）や画像などの非構造化デジタルドキュメントの表形式データは、特に複雑なテーブルの場合、構造とスタイルが複雑で多様であるため、構造化された機械可読形式に解析するのは困難です。このホワイトペーパーでは、正確なテーブル構造認識機能であるSplit、Embed and Merge（SEM）を紹介します。最初の段階では、FCNを使用して、テーブルの行（列）セパレーターの潜在的な領域を予測し、テーブル内の基本グリッドの境界ボックスを取得します。第2段階では、RoIAlignを介して各グリッドに対応する視覚的特徴を抽出するだけでなく、既製の認識機能とBERTを使用して意味的特徴を抽出します。両方の融合された機能は、各テーブルグリッドを特徴づけるために使用されます。各グリッドにセマンティック機能を追加することで、視覚的な観点から見たテーブル構造のあいまいさの問題をある程度解決し、より高い精度を実現できることがわかりました。最後に、これらの基本グリッドのマージを自己回帰方式で処理します。対応するマージ結果は、注意メカニズムの注意マップによって学習されます。提案手法では、複雑なテーブルでもテーブルの構造をよく認識できます。 SEMは、SciTSRデータセットで96.9％の平均Fメジャーを達成できます。これは、他の方法を大幅に上回っています。他の公的に利用可能なテーブル構造認識データセットでの広範な実験は、私たちのモデルが最先端を達成していることを示しています。

The task of table structure recognition is to recognize the internal structure of a table, which is a key step to make machines understand tables. However, tabular data in unstructured digital documents, e.g. Portable Document Format (PDF) and images, are difficult to parse into structured machine-readable format, due to complexity and diversity in their structure and style, especially for complex tables. In this paper, we introduce Split, Embed and Merge (SEM), an accurate table structure recognizer. In the first stage, we use the FCN to predict the potential regions of the table row (column) separators, so as to obtain the bounding boxes of the basic grids in the table. In the second stage, we not only extract the visual features corresponding to each grid through RoIAlign, but also use the off-the-shelf recognizer and the BERT to extract the semantic features. The fused features of both are used to characterize each table grid. We find that by adding additional semantic features to each grid, the ambiguity problem of the table structure from the visual perspective can be solved to a certain extent and achieve higher precision. Finally, we process the merging of these basic grids in a self-regression manner. The correspondent merging results is learned by the attention maps in attention mechanism. With the proposed method, we can recognize the structure of tables well, even for complex tables. SEM can achieve an average F-Measure of 96.9% on the SciTSR dataset which outperforms other methods by a large margin. Extensive experiments on other publicly available table structure recognition datasets show that our model achieves state-of-the-art.

updated: Mon Jul 12 2021 06:26:19 GMT+0000 (UTC)

published: Mon Jul 12 2021 06:26:19 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト