Dark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval Icelandic

Sep 7, 2026, 3:30 PM
30m
Room 3

Room 3

Speaker

Katrín Lísa van der Linde Mikaelsdóttir (University of Iceland)

Description

Iceland’s literary heritage includes around 1000 medieval manuscripts and charters dating from the eleventh century onward, alongside several thousands of post-medieval documents. Although most of the preserved vernacular corpus has already been digitized and catalogued, Icelandic remains a low-resource language for Handwritten Text Recognition. Publicly available models are scarce due to a shortage of ground truth data, compounded by material and palaeographical obstacles. Vellum remained in use well into the seventeenth century, leaving a corpus characterized by darkened or heavily damaged leaves, dense cursive scripts that are heavily abbreviated, and phonologically significant diacritics marked by hairline strokes. This presentation provides a deep-dive analysis of the first public Transkribus model developed for medieval Icelandic. Using the Character Error Rate analysis tool CERatosaurus, I examine model performance across varying material conditions, layout segmentation errors, and specific language- and character-specific confusions. Finally, I present tested workflows that offer practical approaches for building HTR pipelines applicable to underresourced languages facing similar palaeographic and material constraints.

Presentation materials

There are no materials yet.