Matchbox: Creating a Combined Recognition Model for Early Medieval Celtic Languages and Latin

Sep 7, 2026, 1:40 PM
10m
Room 3

Room 3

Speaker

Bernhard Bauer (University of Graz)

Description

The early medieval period in Western Europe was marked by intense linguistic and cultural interaction, as vividly attested in the glossed manuscripts of the era. These glosses are found in the margins or between the lines of Latin texts. They provide unparalleled evidence of the multilingual environment in which manuscripts were produced, copied, and read. The ERC-project GlossIT addresses the technical and scholarly challenges of recognising and transcribing these complex sources by developing text recognition models using the eScriptorium/Kraken platform. Our work focuses on the creation of robust, multilingual recognition models tailored to the unique scripts and language mixtures found in glossed manuscripts. Specifically, we are training models for both Irish minuscules and Carolingian minuscules, drawing on high-quality transcriptions produced within the GlossIT project. The training data encompasses a rich array of languages: Old Irish, Old Breton, Old Welsh, Latin, and Greek, reflecting the true diversity of the glossing tradition on Priscian’s Latin grammar and the Venerable Bede’s computistical works. This multilinguality is not only a technical challenge for ATR but also a scholarly opportunity, as it enables new forms of comparative and contact-linguistic research. The presentation will give an overview on the source material, the workflow and the results of using ATR within GlossIT.

Presentation materials

There are no materials yet.