Multi-script Handwritten Digit Recognition Using Multi-task Learning

Mesay Samuel Gondere; Lars Schmidt-Thieme; Durga Prasad Sharma; Randolf Scholz

マルチタスク学習を使用したマルチスクリプト手書き数字認識

手書き数字認識は、機械学習で広く研究されている分野の1つです。 MNISTデータセットでの手書き数字認識に関する幅広い研究とは別に、さまざまなスクリプト認識に関する他の多くの研究があります。ただし、堅牢で多目的なシステムの開発を促進するマルチスクリプト数字認識ではあまり一般的ではありません。さらに、マルチスクリプトの数字認識に取り組むことで、たとえばスクリプトの分類を関連タスクと見なして、マルチタスク学習が可能になります。マルチタスク学習は、関連するタスクに含まれる情報を使用した誘導転送を通じてモデルのパフォーマンスを向上させることは明らかです。したがって、本研究では、マルチタスク学習を用いたマルチスクリプト手書き数字認識を検討する。問題の解決策を示す具体的な事例として、アムハラ語の手書き文字認識も実験されます。ラテン語、アラビア語、カンナダ語を含む3つのスクリプトの手書き数字を調べて、個々のタスクを再編成したマルチタスクモデルが有望な結果を示していることを示しています。この研究では、個々のタスクの予測を使用する新しい方法が提案され、分類のパフォーマンスを支援し、メインタスクの目的のためにさまざまな損失を正規化しました。この発見は、ベースラインおよび従来のマルチタスク学習モデルを上回っています。さらに重要なことに、マルチタスク学習の課題の1つである、タスクのさまざまな損失に重みを付ける必要がなくなりました。

Handwritten digit recognition is one of the extensively studied area in machine learning. Apart from the wider research on handwritten digit recognition on MNIST dataset, there are many other research works on various script recognition. However, it is not very common for multi-script digit recognition which encourage the development of robust and multipurpose systems. Additionally working on multi-script digit recognition enables multi-task learning, considering the script classification as a related task for instance. It is evident that multi-task learning improves model performance through inductive transfer using the information contained in related tasks. Therefore, in this study multi-script handwritten digit recognition using multi-task learning will be investigated. As a specific case of demonstrating the solution to the problem, Amharic handwritten character recognition will also be experimented. The handwritten digits of three scripts including Latin, Arabic and Kannada are studied to show that multi-task models with reformulation of the individual tasks have shown promising results. In this study a novel way of using the individual tasks predictions was proposed to help classification performance and regularize the different loss for the purpose of the main task. This finding has outperformed the baseline and the conventional multi-task learning models. More importantly, it avoided the need for weighting the different losses of the tasks, which is one of the challenges in multi-task learning.

updated: Tue Jun 15 2021 16:30:37 GMT+0000 (UTC)

published: Tue Jun 15 2021 16:30:37 GMT+0000 (UTC)

arXiv

参考文献 (このサイトで利用可能なもの) / References (only if available on this site)

被参照文献 (このサイトで利用可能なものを新しい順に) / Citations (only if available on this site, in order of most recent)

Amazon.co.jpアソシエイト