EAML: Ensemble Self-Attention-based Mutual Learning Network for Document Image Classification

Souhail Bakkali; Ziheng Ming; Mickael Coustaty; Marçal Rusiñol

EAML: 文書画像分類のためのアンサンブル自己注意ベースの相互学習ネットワーク

最近、複雑なディープニューラルネットワークは、文書画像の分類や文書検索などのさまざまな文書理解タスクにおいて大きな関心を集めています。多くの文書タイプには明確な視覚スタイルがあるため、文書画像を分類するためにディープ CNN を使用して視覚的特徴のみを学習すると、クラス間の区別が低く、カテゴリ間のクラス内構造のばらつきが大きいという問題に直面しました。同時に、特定の文書画像内の対応する視覚的特性と合わせて学習されるテキストレベルの理解により、精度の点で分類パフォーマンスが大幅に向上しました。この論文では、アンサンブル訓練可能なネットワークのブロックとして機能する自己注意ベースの融合モジュールを設計します。これにより、トレーニング段階を通じて画像とテキストのモダリティの判別機能を同時に学習することができます。さらに、トレーニング段階で画像モダリティとテキストモダリティの間でポジティブな知識を伝達することで、相互学習を促進します。この制約は、従来の教師あり設定に新しい正則化項として切り捨てカルバック・ライブラー発散損失 Tr-KLD-Reg を追加することで実現されます。私たちの知る限り、相互学習アプローチと自己注意ベースの融合モジュールを活用して文書画像分類を実行するのはこれが初めてです。実験結果は、シングルモーダルおよびマルチモーダルモダリティの精度の観点から、我々のアプローチの有効性を示しています。したがって、提案されたアンサンブル自己注意ベースの相互学習モデルは、ベンチマーク RVL-CDIP および Tobacco-3482 データセットに基づく最先端の分類結果よりも優れています。

In the recent past, complex deep neural networks have received huge interest in various document understanding tasks such as document image classification and document retrieval. As many document types have a distinct visual style, learning only visual features with deep CNNs to classify document images have encountered the problem of low inter-class discrimination, and high intra-class structural variations between its categories. In parallel, text-level understanding jointly learned with the corresponding visual properties within a given document image has considerably improved the classification performance in terms of accuracy. In this paper, we design a self-attention-based fusion module that serves as a block in our ensemble trainable network. It allows to simultaneously learn the discriminant features of image and text modalities throughout the training stage. Besides, we encourage mutual learning by transferring the positive knowledge between image and text modalities during the training stage. This constraint is realized by adding a truncated-Kullback-Leibler divergence loss Tr-KLD-Reg as a new regularization term, to the conventional supervised setting. To the best of our knowledge, this is the first time to leverage a mutual learning approach along with a self-attention-based fusion module to perform document image classification. The experimental results illustrate the effectiveness of our approach in terms of accuracy for the single-modal and multi-modal modalities. Thus, the proposed ensemble self-attention-based mutual learning model outperforms the state-of-the-art classification results based on the benchmark RVL-CDIP and Tobacco-3482 datasets.

updated: Thu May 11 2023 16:05:03 GMT+0000 (UTC)

published: Thu May 11 2023 16:05:03 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト