Speaker
Description
Iceland’s literary heritage includes around 1000 medieval manuscripts and charters dating from the eleventh century onward, alongside several thousands of post-medieval documents. Although most of the preserved vernacular corpus has already been digitized and catalogued, Icelandic remains a low-resource language for Handwritten Text Recognition. Publicly available models are scarce due to a shortage of ground truth data, compounded by material and palaeographical obstacles. Vellum remained in use well into the seventeenth century, leaving a corpus characterized by darkened or heavily damaged leaves, dense cursive scripts that are heavily abbreviated, and phonologically significant diacritics marked by hairline strokes. This presentation provides a deep-dive analysis of the first public Transkribus model developed for medieval Icelandic. Using the Character Error Rate analysis tool CERatosaurus, I examine model performance across varying material conditions, layout segmentation errors, and specific language- and character-specific confusions. Finally, I present tested workflows that offer practical approaches for building HTR pipelines applicable to underresourced languages facing similar palaeographic and material constraints.