Speaker
Description
Medieval scripts present Handwritten Text Recognition tools with a wide variety of challenges due to highly diverse orthographic features specific to different languages and scripts. Using medieval Greek and Arabic as case studies, this talk examines the handling of features from diacritical marks (accents, breathing marks, vowels, and consonant pointing) to different categories of abbreviation (suspension, contraction, tachyographic, etc). It further examines how modern Unicode encoding -- such as normalization and character-doubling issues -- interacts with these historical script features.
Scholars producing ground truth for these scripts must make choices -- conscious or unconscious -- as to how they will handle the representation of these features. The talk will present several experiments examining the impacts of these different feature categories on HTR model accuracy. As more HTR projects publish and share their ground truth data, we must recognize that inconsistent transcription standards can significantly impact the interoperability of separately published datasets.