Speaker
Description
Medieval manuscripts, unlike modern printed documents, are typically characterized by non-uniform script and complex page layout, both with respect to larger regions (e.g., main text regions and marginal regions) and individual text lines. Especially manuscripts featuring extensive annotations such as glosses or reading signs tend to increase the page layout complexity significantly: this makes the task of automated layout segmentation very difficult. In this presentation, we report on layout analysis experiments comparing fine-tuned Kraken 6.0 BLLA (base-line layout analysis) and YOLO26 nano/small/medium models for the detection of both page regions and text line masks. Training and evaluation are based on ground truth data that was created in the GlossIT project, featuring 874 individual manuscript pages from five different manuscripts (from the 9th and 10th centuries) annotated using 6 region classes and 4 line classes. This dataset allows for a detailed comparison of model performance on different glossed manuscripts.