HTR of Normalized Latin Texts: Insights from Liturgical Manuscripts

Sep 7, 2026, 4:30 PM
20m
Room 2

Room 2

Speaker

Paweł Figurski (Polish Academy of Sciences)

Description

Nearly eleven hundred sacramentaries copied up to c. 1100 — complete codices, fragments, and palimpsests — survive, dispersed across collections worldwide. The scale of this corpus has kept it outside granular analysis, and the tradition is still described through a handful of types established by earlier scholarship. RITUS+, a tool under development, addresses that scale by treating transcription not as an end but as a query.

The pipeline runs from images or IIIF manifests through Kraken-based recognition to automatic liturgical indexing. Colour analysis separates red from black within each line, so that rubrics are recovered as structure rather than merely as text. Recognition output is then tokenized and matched — by iterative fuzzy search with progressively relaxed thresholds, refined by exact Levenshtein distance — against a curated lexicon of prayers. Each identified formula receives an identifier pointing to the standardized text in the database, not to the manuscript's own spelling.

This normalization is deliberate, and it carries a cost: the tool does not address philological questions about individual variants. What it buys is that an imperfect transcription remains a serviceable query, since for a genre this formulaic the operative unit of recognition is the formula rather than the character. Indexing then permits comparison on content and on sequence, filtering a manuscript against the better-known indexed traditions and isolating what belongs to none of them — the material still awaiting scholarly attention.

Presentation materials

There are no materials yet.