Speaker
Description
Transcription levels, the identification of scribal hands, and workflow architectures are usually discussed in three separate conversations, one editorial, one palaeographical, one technical. I fold them into two decisions taken before any computational result exists: at what grain the written surface is read (allographetic, graphematic, normalised on the text side; glyph, patch, line, page on the image side), and when each reading is committed, early at explicit interfaces or late inside an integrated model. Each decision leaves something behind. Four experiments supply the evidence. In HORAE, liturgical texts are identified on noisy, unregularised ATR output: the regularised text becomes a by-product of retrieval, and the residue (variants, accessory prayers) is the scholarly signal. On the same corpus, sequential, hybrid and integrated architectures for structural annotation reveal a chiasm: sequential workflows commit their projections early, at inspectable interfaces, and suffer from them; integrated ones escape premature projection but hide their commitments, gaining most at the granularities the pipeline had already projected away. The results of the FalsID competition on falsification and imitation detection in writer identification, and a page-level versus patch-level comparison of script and hand modelling confirm the pattern. Against the emerging consensus of a graphematic pivot, I argue the record is not a level but the aligned bundle, of which every level is a view. The choice before us is not between transcription levels, nor even between workflows: it is between architectures that remember their interpretations and architectures that forget them. What remains is then twofold: the residue each projection silences, and the record that outlives our pipelines. Editorial decisions are those that keep the former alive inside the latter.