Speaker
Description
Automatic Text Recognition (ATR) has made manuscript scholarship scalable, but it also intro- duces profound epistemic challenges stemming from the conflation of transcription (the reporting of the visible source text) and edition (the act of scholarly interpretation). We argue that some of the current ATR practices may produce silent forgeries—plausible yet undocumented inferences that threaten the reliability of large-scale textual corpora. Unlike conventional ATR errors, these inferences are not visually grounded and therefore difficult to detect. Furthermore, standard evaluation metrics such as Character Error Rate (CER) are inadequate because they are biased by abbreviation expansion, favoring models that produce transcriptions with expanded abbreviations. To create more reliable, automatically produced textual resources, we propose a layered pipeline framework that enforces a strict separation between three distinct stages: Graphemic Transcription (the visual layer), Pre-Editorial Normalization (the linguistic interpretation layer), and Scholarly Editing (the contextual layer). This approach restores scholarly accountability by making every stage of textual establishment transparent, ensuring that computational outputs become durable, reusable research infrastructures rather than single-use records.
Zoom link for the keynote: https://oeaw-ac-at.zoom.us/j/65915000613?pwd=bmH0pagzGK6QTFt9eqxTAse0HlmETV.1