-
Ana Mihaljević (Institute for the Croatian language)9/7/26, 1:30 PM
The digitization of the Dictionary of the Croatian redaction of Church Slavonic, carried out within the DigiSTIN project, has presented us with a particular challenge: how to efficiently OCR and HTR materials that combine multiple languages and writing systems. The dictionary covers Croatian Church Slavonic, Croatian, English, Latin, and Greek, and uses Latin, Greek, Glagolitic, and Old...
Go to contribution page -
Bernhard Bauer (University of Graz)9/7/26, 1:40 PM
The early medieval period in Western Europe was marked by intense linguistic and cultural interaction, as vividly attested in the glossed manuscripts of the era. These glosses are found in the margins or between the lines of Latin texts. They provide unparalleled evidence of the multilingual environment in which manuscripts were produced, copied, and read. The ERC-project GlossIT addresses the...
Go to contribution page -
Baptiste Queuche (Calfa)9/7/26, 1:50 PM
Calfa is developing general and custom AI models for researchers and cultural heritage professionals, dedicated to recognize the text and analyze documents in non-western languages, such as Arabic, Armenian, Georgian, Syriac, Chinese, Greek. Calfa also offers many open-source datasets for Arabic or Armenian languages.
We recently developed a promising hybrid approach, combining our specific...
Go to contribution page -
Seth Kulick (Linguistic Data Consortium, University of Pennsylvania)9/7/26, 2:00 PM
Historical treebanks — corpora annotated with part-of-speech and syntactic information — are essential resources for linguists studying language change, but they are labor-intensive to build. Existing historical treebanks are consequently small, typically around 1-2 million words, which limits the scope of research they can support. This talk gives an overview of work that explores whether...
Go to contribution page -
Andrew Janco (Princeton University)9/7/26, 2:10 PM
Transformer-based language models have a fixed vocabulary of tokens used to represent words and word parts. Which languages can and cannot be represented with current model vocabularies? When native support does not exist, what are the possibilities and drawbacks of byte pair encoding? This short talk presents a tool to check how well current model vocabularies support the language(s) you’re...
Go to contribution page -
Isabelle Marthot-Santaniello (University of Basel)9/7/26, 2:20 PM
Greek has been used in Egypt from the time of Alexander the Great to beyond the Islamic conquest. During this millennium, Greek, like any language, evolved. It was also in contact with other languages (Demotic Egyptian, Latin, Coptic, Arabic…), a multilingual situation that encouraged the emergence of loanwords and the modification of some formulas. Besides, a wide range of text types were...
Go to contribution page -
Katrín Lísa van der Linde Mikaelsdóttir (University of Iceland)9/7/26, 3:30 PM
Iceland’s literary heritage includes around 1000 medieval manuscripts and charters dating from the eleventh century onward, alongside several thousands of post-medieval documents. Although most of the preserved vernacular corpus has already been digitized and catalogued, Icelandic remains a low-resource language for Handwritten Text Recognition. Publicly available models are scarce due to a...
Go to contribution page -
Christine Roughan (Princeton University)9/7/26, 4:00 PM
Medieval scripts present Handwritten Text Recognition tools with a wide variety of challenges due to highly diverse orthographic features specific to different languages and scripts. Using medieval Greek and Arabic as case studies, this talk examines the handling of features from diacritical marks (accents, breathing marks, vowels, and consonant pointing) to different categories of...
Go to contribution page -
Ephrem Aboud Ishac (Institute for Medieval Research, Austrian Academy of Sciences)9/7/26, 4:30 PM
As Automated Text Recognition (HTR) becomes increasingly integral to Digital Humanities, under-resourced scripts such as Syriac continue to present unique methodological challenges. This presentation outlines an experimental workflow designed to bridge this gap by not only recognizing Syriac handwriting but actively "giving voice" to the texts. The proposed pipeline begins with the HTR...
Go to contribution page
Choose timezone
Your profile timezone: