SCOOP: Source Codes of the Past

Back to the programme

Anagnostes - Towards a Transformer-based OCR System for Ancient Greek Papyri

Elena Chepel & Anton Repushko ・ University of Vienna

Monday, 7th September ・ 13:30 - 15:00 ・ WG1-2 ・ Training as a continuous process ・ Room 1

In this paper, we present a new text recognition tool, Anagnostes, which combines modern approaches in machine learning with optical character recognition (OCR). It has been trained to read Greek literary and documentary papyri from images, with a particular emphasis on recognising not only neat and well-preserved fragments but also cursive and damaged texts. Currently, the model achieves an average character error rate (CER) of 10% across all types of Greek papyri. In our presentation, we will provide a more detailed breakdown of CER across various categories, including literary and documentary texts, as well as clear, abraded, faded, and lacunose manuscripts. We will also discuss a range of challenging features found in papyri, such as highly cursive hands, quadrilinear scripts, heavily abbreviated texts, and the use of special characters in documentary texts. These features are currently covered only sporadically by our model. We will propose future training strategies aimed at overcoming these challenges. Finally, we will discuss how training can be adapted to target other languages represented on papyri, such as Egyptian Hieratic, Egyptian Demotic, Coptic, and Arabic.