2ⁿᵈ Exchange Meeting of SCOOP: International Network for Automated Text Recognition of Historical Sources
University of Vienna
2nd Exchange Meeting, Vienna, September 7-9. 2026
SCOOP (Source Codes of the Past) is an international network dedicated to the automatic transcription and analysis of handwritten historical sources. It brings together people from very different corners of the scholarly world: historians, philologists, and palaeographers, archivists and librarians; computer scientists and machine learning researchers; software engineers asf. What unites them is a shared challenge: how to use automatic/handwritten text recognition (ATR/HTR) to unlock the vast written heritage of the past, and how to do it well. (Read more about SCOOP.)
After a first smaller meeting at Princeton in June of 2025 (for more information see https://fsp-text-edition-blog.univie.ac.at/?p=128), the second SCOOP Exchange Meeting is taking place in Vienna on 7–9 September 2026 hosted by the Institute for Medieval Research at the Austrian Academy of Sciences and the Faculty for Cultural Historical Studies at the University of Vienna.
Over three days, it will bring together around a hundred members of the network for keynotes, parallel working group sessions, round tables, demonstrations, and open discussion formats, along with: as well as an internal working day devoted to the future of SCOOP:
September 7–8: Part of the conference open to registered external visitors
(Keynotes will be streamed via Zoom, see Keynote 1 and Keynote 2.)
September 9: A dedicated internal working session for SCOOP members
September 10 (Associated Event): The OCR/HTR Workshop for Under-represented and Under-resourced Languages organized by Alíz Horváth - Zoom attendance is still possible, please contact Alíz Horváth for the link.(HorvathA@ceu.edu). See the programme.
-
-
8:00 AM
Registration
-
Plenary: Welcome Room 1
Room 1
-
Keynote Room 1
Room 1
-
1
The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text Recognition
Automatic Text Recognition (ATR) has made manuscript scholarship scalable, but it also intro- duces profound epistemic challenges stemming from the conflation of transcription (the reporting of the visible source text) and edition (the act of scholarly interpretation). We argue that some of the current ATR practices may produce silent forgeries—plausible yet undocumented inferences that threaten the reliability of large-scale textual corpora. Unlike conventional ATR errors, these inferences are not visually grounded and therefore difficult to detect. Furthermore, standard evaluation metrics such as Character Error Rate (CER) are inadequate because they are biased by abbreviation expansion, favoring models that produce transcriptions with expanded abbreviations. To create more reliable, automatically produced textual resources, we propose a layered pipeline framework that enforces a strict separation between three distinct stages: Graphemic Transcription (the visual layer), Pre-Editorial Normalization (the linguistic interpretation layer), and Scholarly Editing (the contextual layer). This approach restores scholarly accountability by making every stage of textual establishment transparent, ensuring that computational outputs become durable, reusable research infrastructures rather than single-use records.
Zoom link for the keynote: https://oeaw-ac-at.zoom.us/j/65915000613?pwd=bmH0pagzGK6QTFt9eqxTAse0HlmETV.1
Speaker: Thibault Clérice (ALMAnaCH, Inria)
-
1
-
10:30 AM
Coffee break
-
Working Group 1: WG1-1 - Technologies & Architectures Room 1
Room 1
-
2
Finetuning VLMs in a World of Zero-Shot Frontier Models: A Paradigm Shift toward Knowledge Distillation and Local Deployment
The rapid evolution of frontier Vision-Language Models (VLMs), such as the Gemini 3.5 family, has introduced unprecedented zero-shot capabilities in complex multimodal tasks, including high-fidelity document transcription and precise spatial reasoning via bounding box generation. As these massive, proprietary models increasingly solve general visual-linguistic tasks without task-specific training, the traditional role of fine-tuning is fundamentally shifting. This paper explores the contemporary landscape of VLM fine-tuning, showing that its primary utility has transitioned from baseline task adaptation to targeted knowledge distillation. We outline a framework for leveraging the advanced zero-shot inferences of frontier models as high-quality, synthetic supervisory signals to fine-tune smaller, open-weight models. By treating frontier VLMs as automated annotators or "teachers," we demonstrate how their generalized intelligence can be distilled into compact "student" models. This approach not only dramatically reduces inference latency and operational costs but also addresses strict data privacy and security requirements by enabling robust, offline deployment.
Speaker: William Mattingly (Yale University) -
3
AnandaSky
This presentation introduces AnandaSky, a compact vision-language model developed for line-level transcription of historical sinographic documents, and reports ongoing work toward a new page-level version designed for broader cross-domain use. AnandaSky combines a shallow, high-resolution visual encoder with a Qwen3-0.6B autoregressive decoder. Global visual attention, 10-pixel patches, an uncompressed visual-token prefix, and variable-length attention preserve fine stroke information while keeping the full model to approximately 626 million parameters. Training uses more than four million annotated lines, representing 66.6 million character instances from documents produced between the eighth and twentieth centuries.
AnandaSky demonstrates that compact, domain-specialized vision-language models can outperform much larger general-purpose systems. It achieves character error rates below 1% on five of eight in-domain and held-out evaluation sets and establishes a new state of the art on MTHv2 with 0.92% CER, compared with the previously reported 2.11%. The model also trans fers effectively to unseen document types, including family records and Taoist texts. Error analysis further suggests that, at this level of accuracy, aggregate CER increasingly reflects annotation noise, normalization mismatches, and ambiguous reference transcriptions rather than recognition failures alone.
Building on these findings, the forthcoming version of AnandaSky extends the task from isolated line transcription to integrated page-level document processing. The new model integrates layout analysis and transcription so that complete pages can be processed without assuming perfectly extracted lines. It is trained with a larger and more diverse collection intended to broaden coverage of documentary genres, visual conditions, geographic traditions, and transcription practices. This expansion directly targets the domain shifts observed when models encounter collections whose writing conventions, language, or page organization differ from the training distribution. Preliminary experiments suggest that this page-level approach improves robustness across collections while preserving the strong transcription performance of the original model, advancing AnandaSky toward an end-to-end framework for scalable historical document digitization.
Speaker: Colin Brisson (Ecolé pratique des hautes études) -
4
Metatools & Technological Agnosticism
It seems the age of thinking in tools is at an end. While Large Language Models offer rapid text processing, generic chat interfaces struggle with complex layout analysis, historical scripts, and verifiable ground truth, and especially control over workflows. As a toolbox, Transkribus follows a different path, aiming for a necessary alternative and an essential complement to LLM-driven workflows. As a technologically agnostic metatool, historical document analysis is decoupled from opaque black-box architectures by utilizing open standards like PAGE XML for precise layout segmentation and deterministic ATR. Reproducible, fine-tuned models with downstream LLMs create a (more) transparent, hybrid pipeline that mitigates hallucinations, ethical, and ecological concerns while scaling digital humanities research.
Speaker: Andy Stauder (READ COOP)
-
2
-
12:30 PM
Lunch
-
Working Group 4: Language Challenges I Room 3
Room 3
-
5
Four Scripts, Five Languages, One Digitalization Challenge
The digitization of the Dictionary of the Croatian redaction of Church Slavonic, carried out within the DigiSTIN project, has presented us with a particular challenge: how to efficiently OCR and HTR materials that combine multiple languages and writing systems. The dictionary covers Croatian Church Slavonic, Croatian, English, Latin, and Greek, and uses Latin, Greek, Glagolitic, and Old Cyrillic scripts. While the more recent volumes are available in digital form, the older volumes exist only in print, making them an ideal testing ground for different approaches to automatic text recognition. The complexity of the project is further increased by the dictionary's underlying corpus: several million paper slips written in more than twenty different handwrittings, containing texts in Latin, Greek, and Old Cyrillic. This manuscript material is itself currently undergoing digitization.
In this short talk, we will share our experience of developing the dictionary's online portal and digitizing both its printed and manuscript materials. Particular attention will be given to the advantages and limitations of OCR, HTR, custom-trained models, and LLM-based approaches when working with multilingual and multiscript historical materials.
Speaker: Ana Mihaljević (Institute for the Croatian language) -
6
Matchbox: Creating a Combined Recognition Model for Early Medieval Celtic Languages and Latin
The early medieval period in Western Europe was marked by intense linguistic and cultural interaction, as vividly attested in the glossed manuscripts of the era. These glosses are found in the margins or between the lines of Latin texts. They provide unparalleled evidence of the multilingual environment in which manuscripts were produced, copied, and read. The ERC-project GlossIT addresses the technical and scholarly challenges of recognising and transcribing these complex sources by developing text recognition models using the eScriptorium/Kraken platform. Our work focuses on the creation of robust, multilingual recognition models tailored to the unique scripts and language mixtures found in glossed manuscripts. Specifically, we are training models for both Irish minuscules and Carolingian minuscules, drawing on high-quality transcriptions produced within the GlossIT project. The training data encompasses a rich array of languages: Old Irish, Old Breton, Old Welsh, Latin, and Greek, reflecting the true diversity of the glossing tradition on Priscian’s Latin grammar and the Venerable Bede’s computistical works. This multilinguality is not only a technical challenge for ATR but also a scholarly opportunity, as it enables new forms of comparative and contact-linguistic research. The presentation will give an overview on the source material, the workflow and the results of using ATR within GlossIT.
Speaker: Bernhard Bauer (University of Graz) -
7
How hybrid HTR+VLM approaches are unlocking the processing of under-resourced, non-western languages
Calfa is developing general and custom AI models for researchers and cultural heritage professionals, dedicated to recognize the text and analyze documents in non-western languages, such as Arabic, Armenian, Georgian, Syriac, Chinese, Greek. Calfa also offers many open-source datasets for Arabic or Armenian languages.
We recently developed a promising hybrid approach, combining our specific OCR/HTR models with customized VLMs.
This approach allows for the processing of a very wide variety of languages and scripts with excellent accuracy rates on both handwritten documents and printed archives.
Researchers can now obtain a specialized AI solution tailored to their corpus, enabling not only OCR/HTR processing but also analysis and understanding of the documents layout according to their project specifications.
Several tools are offered free of charge online or for download by Calfa so that researchers can try these solutions on their corpus.Speaker: Baptiste Queuche (Calfa) -
8
Using OCR to expand historical treebanks
Historical treebanks — corpora annotated with part-of-speech and syntactic information — are essential resources for linguists studying language change, but they are labor-intensive to build. Existing historical treebanks are consequently small, typically around 1-2 million words, which limits the scope of research they can support. This talk gives an overview of work that explores whether that bottleneck can be broken by applying automatic syntactic annotation to large amounts of historical text containing OCR or OCR-like errors, dramatically increasing treebank size without manual annotation. The central question is whether automatic annotation on such imperfect historical text is accurate enough to serve as a reliable basis for modeling syntactic change. Current work focuses on the case of Yiddish texts from the 16th and 17th centuries.
Speaker: Seth Kulick (Linguistic Data Consortium, University of Pennsylvania) -
9
Which Languages Are In-Vocabulary?
Transformer-based language models have a fixed vocabulary of tokens used to represent words and word parts. Which languages can and cannot be represented with current model vocabularies? When native support does not exist, what are the possibilities and drawbacks of byte pair encoding? This short talk presents a tool to check how well current model vocabularies support the language(s) you’re working with. It joins data for more than 8,000 languages from the Unicode CLDR and Glottolog, with common vocabularies in tiktoken and HuggingFace AutoTokenizer.
Speaker: Andrew Janco (Princeton University) -
10
The language of Greek papyri? One millennium of writing and its complexity
Greek has been used in Egypt from the time of Alexander the Great to beyond the Islamic conquest. During this millennium, Greek, like any language, evolved. It was also in contact with other languages (Demotic Egyptian, Latin, Coptic, Arabic…), a multilingual situation that encouraged the emergence of loanwords and the modification of some formulas. Besides, a wide range of text types were written on papyri and ostraca, from deluxe editions of literature to text receipts, accountings and private letters, using various writing styles. How can these shades of Greek be taken into account in a computer-assisted exploration of this material?
Speaker: Isabelle Marthot-Santaniello (University of Basel)
-
5
-
Working Group 1: Training as a continuous process Room 1
Room 1
-
11
Piloting HTR in the Library: Experiments and Groundwork
Since 2023, members of Penn Libraries’ Cultural Heritage Computing, Schoenberg Institute for Manuscript Studies, and Research Data and Digital Scholarship groups have conducted workshops and pilot projects to explore workflows for handwritten text recognition (HTR) and its potential to support transcription of library collections for discovery, access, and computational use. This talk will focus on our exploratory work with HTR models and workflows as groundwork for expanding access to Penn Libraries’ diverse manuscript collections.
Speaker: Doug Emery (University of Pennsylvania Libraries) -
12
Multi-Model OCR Arbitration: High-Precision OCR for 19th Century Administrative Mass Sources
The transcription of historical administrative sources presents a distinctive set of challenges for automated text recognition: dense, formulaic layouts, highly abbreviated Latin and German terminology, and the cumulative degradation typical of large serial document corpora. This paper presents a high-precision OCR pipeline developed within the Unlocking the Schematismus project, designed to meet the specific demands of the Austrian Schematismus series, annual administrative directories spanning the Habsburg monarchy from the late eighteenth to the early twentieth century.
The pipeline integrates three complementary transcription models, each fine-tuned on annotated material from the corpus: a Kraken instance adapted to its typographic conventions, and two transformer-based vision models, Florence and TrOCR. Rather than selecting a single model in advance or averaging across outputs, the architecture adds a fourth component: a vision-language model used as a selection stage, which is shown the three candidate transcriptions together with the original image snippet and chooses the most plausible reading. Because this stage conditions on the image as well as on the textual candidates, it can resolve disagreements that neither confidence-based voting nor purely textual reranking would settle correctly.
The resulting pipeline achieves a character error rate of 0.2%, positioning it among the most precise HTR/OCR systems reported for comparable historical document types. Beyond the technical result, the paper reflects on the broader methodological implications: what does near-zero error mean in a corpus where orthographic inconsistency and abbreviation are themselves historically significant? And to what extent does the selection model encode implicit editorial decisions that warrant explicit philological scrutiny?
The Schematismus pipeline thus raises questions familiar from other corners of the digital humanities: the relationship between computational efficiency and scholarly responsibility, and the degree to which automated systems can, or should, replicate the interpretive judgments of the human editor.
Speakers: Wolfgang Göderle (University of Graz | University of Passau | MPI GEA), David Fleischhacker (University of Graz) -
13
From PDF to mARkdown: A Scalable Arabic OCR Pipeline for expanding the OpenITI Corpus
This talk presents a scalable pipeline for converting digitised Arabic books from PDFs into structured mARkdown for integration into the OpenITI corpus. The workflow begins by preparing the PDFs and processing their opening pages to extract the bibliographic metadata required to identify works and editions. The documents are then transcribed in parallel by several complementary recognition systems that use different approaches to document OCR and vision-language processing. Each system’s output is converted into a common page-based representation that records its text, lines, structural labels, coordinates and recognition confidence when available. The pipeline aligns the resulting tokens and compares competing readings using weighted agreement across the systems. The selected readings are combined with the available layout and document-structure information to produce a page-aware consensus containing headings, paragraphs, line breaks, footnotes, and the division of poetry into verses and hemistichs. Although currently optimised for printed books, the pipeline also provides an experimental workflow for manuscripts, which require more extensive manual correction.
The resulting consensus and extracted metadata are presented in an interactive review environment in which users can compare alternative readings and correct the bibliographic information, transcription and layout. Reviewed documents are exported both as corpus-ready OpenITI mARkdown and as ground-truth data for evaluation and for fine-tuning locally trainable OCR models.
The pipeline is designed to reduce the time required for OCR and scholarly revision while keeping processing costs limited. Its modular architecture allows individual recognition and processing components to be enabled, disabled or replaced. The talk will describe the workflow from PDF ingestion to final mARkdown export and compare the textual and structural accuracy of the individual recognition approaches with that of the resulting consensus, alongside measurements of processing time and cost. Finally, it will consider how complementary recognition, structured consensus and reusable human correction can support the large-scale, quality-controlled expansion of OpenITI.
Speaker: Alicia González Martínez (Hamburg University) -
14
Anagnostes - Towards a Transformer-based OCR System for Ancient Greek Papyri
In this paper, we present a new text recognition tool, Anagnostes, which
combines modern approaches in machine learning with optical character
recognition (OCR). It has been trained to read Greek literary and
documentary papyri from images, with a particular emphasis on
recognising not only neat and well-preserved fragments but also cursive
and damaged texts. Currently, the model achieves an average character
error rate (CER) of 10% across all types of Greek papyri. In our
presentation, we will provide a more detailed breakdown of CER across
various categories, including literary and documentary texts, as well as
clear, abraded, faded, and lacunose manuscripts. We will also discuss a
range of challenging features found in papyri, such as highly cursive
hands, quadrilinear scripts, heavily abbreviated texts, and the use of
special characters in documentary texts. These features are currently
covered only sporadically by our model. We will propose future training
strategies aimed at overcoming these challenges. Finally, we will
discuss how training can be adapted to target other languages
represented on papyri, such as Egyptian Hieratic, Egyptian Demotic,
Coptic, and Arabic.Speakers: Elena Chepel (University of Vienna), Anton Repushko (Independent Researcher)
-
11
-
Working Group 3: Transcription approaches I (Palaeography in focus) Room 2
Room 2
-
15
Computational Paleography through Automatic Text Recognition
Computational methods have transformed many areas of historical research, yet
paleography has seen comparatively few algorithms and tools that can operate at
scale while remaining adaptable to scholarly practices across different
scripts. This is not due to a lack of interest in digital methods. The growth
of digital repositories, annotation platforms such as DigiPal, and numerous
prototypes using modern machine learning all demonstrate the field’s
openness to computational approaches. What remains lacking are methods
that can support paleographic investigation without narrowing it to a
predefined set of categories or research questions.This presentation explores automatic text recognition (ATR) as one such
method. Rather than treating ATR only as a means of producing
transcriptions, it considers how recognition models can serve as
instruments for paleographic research. Their combination of
adaptability, high throughput, and introspectability makes it possible
to examine large bodies of material while retaining access to the
individual written forms on which an analysis is based.Placed near the bottom of the ladder of abstraction, this approach
offers a way to formulate research questions and validate observations
without presupposing an established paleographic framework. It may
therefore be particularly valuable for the study of non-Western and
minority writing traditions, for which inherited classifications may be
incomplete, inappropriate, or altogether absent.Speaker: Benjamin Kiessling (ALMAnaCH, Inria Paris) -
16
Towards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic Manuscripts
Arabic manuscript traditions present a constellation of challenges that continue to resist the assumptions underlying most contemporary Handwritten Text Recognition (HTR) systems. Unlike Latin scripts, Arabic is written right-to-left in a fully cursive hand in which most, but not all, letters are connected to their neighbors, and the graphical form of each letter changes depending on its position within the word—initial, medial, final, or isolated. This positional morphology means that no letter has a single canonical shape; its geometry is always a function of its immediate lexical context. Compounding this, Arabic manuscript pages rarely exhibit a consistent baseline: words within a single line float at varying angles and heights, displaced by diacritics, sublinear descenders, and the calligrapher's deliberate compositional choices. At the level of scribal tradition, the challenge is further complicated—the same word written in Naskh, Thuluth, Maghrebi, or Ruq'a hands may share almost no visual geometry, meaning that any robust recognition system must account for inter-style variation before it can meaningfully generalize across corpora. These compounding factors make character-level and line-level recognition paradigms ill-suited as primary analytical units for Arabic manuscripts. We argue that the word is the natural unit of analysis, as it is the smallest graphically self-contained entity that encodes both lexical identity and scribal style simultaneously. In this paper, we present a word-level graphometric framework that addresses these challenges through geometric vectorization: each word image is binarized, skeletonized into a stroke graph, and characterized by a feature vector encoding structural complexity, stroke-width variation, path geometry, and ink density. We describe a quality-aware cluster profiling system that builds statistical fingerprints for each word type across manuscript instances, and a geometry-only classifier that assigns new, untranscribed word images to existing clusters without relying on prior lexical knowledge. We present the results of experiments conducted on a segmented Arabic manuscript corpus, demonstrating that computational paleography—the systematic, quantitative study of scribal hands—can expose the within-corpus variation that HTR models must learn to navigate, and that the graphometric profiles produced by our framework offer a principled basis for building training corpora that are representative of scribal diversity.
Speaker: Sajjad Nikfahm-Khubravan (Roshan Institute for Persian Studies, University of Maryland) -
17
Leveraging HTR for Script Characterization: Toward a Unified Computational Framework for Palaeographical Analysis
The analysis of medieval handwriting lies at the heart of palaeographical research. While Automatic Text Recognition (ATR) systems are primarily developed to produce transcriptions, they also generate rich data that can be leveraged to study the scripts themselves. This presentation will explore how ATR data can be exploited to develop new computational tools for palaeographical analysis. Starting from text-line images and their transcriptions, the proposed framework automatically detects and models character instances without requiring manual character-level annotation. It produces visual prototypes representing average letterforms, together with metrological descriptors encoding their dimensions and spatial positions within the text line. These representations provide quantitative descriptions of handwriting while remaining readily interpretable by palaeographers. The framework will be illustrated through several case studies in Latin palaeography, focusing on book scripts and covering different levels of analysis, including script-type comparison, scribal identification, intra-scribal handwriting variability, and the study of scribal digraphism.
Speaker: Malamatenia Vlachou (IRHT/CNRS-ENPC) -
18
Workflows and granularities: text, image, and what remains
Transcription levels, the identification of scribal hands, and workflow architectures are usually discussed in three separate conversations, one editorial, one palaeographical, one technical. I fold them into two decisions taken before any computational result exists: at what grain the written surface is read (allographetic, graphematic, normalised on the text side; glyph, patch, line, page on the image side), and when each reading is committed, early at explicit interfaces or late inside an integrated model. Each decision leaves something behind. Four experiments supply the evidence. In HORAE, liturgical texts are identified on noisy, unregularised ATR output: the regularised text becomes a by-product of retrieval, and the residue (variants, accessory prayers) is the scholarly signal. On the same corpus, sequential, hybrid and integrated architectures for structural annotation reveal a chiasm: sequential workflows commit their projections early, at inspectable interfaces, and suffer from them; integrated ones escape premature projection but hide their commitments, gaining most at the granularities the pipeline had already projected away. The results of the FalsID competition on falsification and imitation detection in writer identification, and a page-level versus patch-level comparison of script and hand modelling confirm the pattern. Against the emerging consensus of a graphematic pivot, I argue the record is not a level but the aligned bundle, of which every level is a view. The choice before us is not between transcription levels, nor even between workflows: it is between architectures that remember their interpretations and architectures that forget them. What remains is then twofold: the residue each projection silences, and the record that outlives our pipelines. Editorial decisions are those that keep the former alive inside the latter.
Speaker: Dominique Stutzmann (IRHT-CNRS / HU Berlin)
-
15
-
3:00 PM
Coffee break
-
Working Group 4: Language Challenges II Room 3
Room 3
-
19
Dark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval Icelandic
Iceland’s literary heritage includes around 1000 medieval manuscripts and charters dating from the eleventh century onward, alongside several thousands of post-medieval documents. Although most of the preserved vernacular corpus has already been digitized and catalogued, Icelandic remains a low-resource language for Handwritten Text Recognition. Publicly available models are scarce due to a shortage of ground truth data, compounded by material and palaeographical obstacles. Vellum remained in use well into the seventeenth century, leaving a corpus characterized by darkened or heavily damaged leaves, dense cursive scripts that are heavily abbreviated, and phonologically significant diacritics marked by hairline strokes. This presentation provides a deep-dive analysis of the first public Transkribus model developed for medieval Icelandic. Using the Character Error Rate analysis tool CERatosaurus, I examine model performance across varying material conditions, layout segmentation errors, and specific language- and character-specific confusions. Finally, I present tested workflows that offer practical approaches for building HTR pipelines applicable to underresourced languages facing similar palaeographic and material constraints.
Speaker: Katrín Lísa van der Linde Mikaelsdóttir (University of Iceland) -
20
Impacts of Script Features on Text Recognition: Experiments with Greek and Arabic
Medieval scripts present Handwritten Text Recognition tools with a wide variety of challenges due to highly diverse orthographic features specific to different languages and scripts. Using medieval Greek and Arabic as case studies, this talk examines the handling of features from diacritical marks (accents, breathing marks, vowels, and consonant pointing) to different categories of abbreviation (suspension, contraction, tachyographic, etc). It further examines how modern Unicode encoding -- such as normalization and character-doubling issues -- interacts with these historical script features.
Scholars producing ground truth for these scripts must make choices -- conscious or unconscious -- as to how they will handle the representation of these features. The talk will present several experiments examining the impacts of these different feature categories on HTR model accuracy. As more HTR projects publish and share their ground truth data, we must recognize that inconsistent transcription standards can significantly impact the interoperability of separately published datasets.
Speaker: Christine Roughan (Princeton University) -
21
Syriac Manuscripts That Can Talk: Giving Voice to Under-Resourced Scripts in HTR
As Automated Text Recognition (HTR) becomes increasingly integral to Digital Humanities, under-resourced scripts such as Syriac continue to present unique methodological challenges. This presentation outlines an experimental workflow designed to bridge this gap by not only recognizing Syriac handwriting but actively "giving voice" to the texts. The proposed pipeline begins with the HTR processing of Syriac manuscripts using Transkribus, followed by the automated transliteration of the recognized text into Latin script. Once transliterated, Text-to-Speech (TTS) technologies are applied to generate audio output. Crucially, by combining Transkribus-based HTR, transliteration, and TTS, this workflow aims to make historical Syriac sources accessible to individuals with visual impairments, allowing them to hear texts that were previously out of reach. Furthermore, it offers new avenues for linguistic analysis and broader public dissemination.
Speaker: Ephrem Aboud Ishac (Institute for Medieval Research, Austrian Academy of Sciences)
-
19
-
Working Group 1: Experimental approaches Room 1
Room 1
-
22
Utilization of the Open WebUI platform in Arkindex
Arkindex (a document processing platform developed by TEKLIA) represents a promising open-source alternative to Transcribus and e-Scriptorium. Based on modular approach Arkindex implements machine learning algorithms that can be applied for layout analysis, transcription, structuration, named entity recognition, and information extraction of text documents as well as object segmentation, visual description, and metadata extraction of photographs and visual objects. The main aim of the current project is to expand the capabilities of Arkindex by the implementation of suitable open-weight LLMs and VLMs for ATR/HTR tasks managed within local installation of Open WebUI platform. The suggested solution involves the development of specialized “Worker” (a Docker image) able to utilize the compute power of these models via API endpoint provided by the instance of Open WebUI platform.
Speaker: Michal Racyn (Masaryk University) -
23
Document Layout Analysis for Glossed Medieval Manuscripts: Comparing Kraken 6 and YOLO26
Medieval manuscripts, unlike modern printed documents, are typically characterized by non-uniform script and complex page layout, both with respect to larger regions (e.g., main text regions and marginal regions) and individual text lines. Especially manuscripts featuring extensive annotations such as glosses or reading signs tend to increase the page layout complexity significantly: this makes the task of automated layout segmentation very difficult. In this presentation, we report on layout analysis experiments comparing fine-tuned Kraken 6.0 BLLA (base-line layout analysis) and YOLO26 nano/small/medium models for the detection of both page regions and text line masks. Training and evaluation are based on ground truth data that was created in the GlossIT project, featuring 874 individual manuscript pages from five different manuscripts (from the 9th and 10th centuries) annotated using 6 region classes and 4 line classes. This dataset allows for a detailed comparison of model performance on different glossed manuscripts.
Speaker: Tristan Repolusk (Department of Digital Humanities, University of Graz) -
24
Finding our way through the National Archives: HTR for indexing a large corpus
The English common law rests on a documentary record of staggering size and frustrating inaccessibility. At least eleven million pages of medieval legal manuscripts survive at the National Archives at Kew, the product of a documentary machine at Westminster and county courts that churned out records at a rate of tens of thousands per year. For comparison, roughly six to seven thousand literary manuscripts survive from the same period and place. These records are digitized and freely available online, but this has not rendered them accessible. The indices that exist are an array of finding aids and researcher-produced spreadsheets that are not connected and do not cover the full archive, and the vast majority of the manuscripts remain untranscribed. Thus, our project asks not how faithfully we might reproduce each page, but how much of this vast archive can be made navigable as quickly and simply as possible to the historians who need it. We built an open-source pipeline aimed at rough, mostly-correct transcription at scale. Trained on roughly 4,000 hand-transcribed lines, it pairs a neural line-segmentation with a convolutional–recurrent network (a CNN feeding an LSTM, decoded with CTC and a Latin n-gram language model); this brings word error down to around twelve percent and character error to around four, accurate enough to let a researcher locate, sort, and triage cases that would otherwise cost a trained paleographer some five minutes a line to decipher. We paired this with a search engine and a Claude-based assistant to sort through HTR errors and variations to return all relevant results. The tool is designed to work in tandem with an experienced reader, not to replace one; its errors are, by design, the kind a human can readily catch. This paper describes the project from the historian’s side as the practical work of turning the parchment maze of the National Archives into a searchable record of medieval life.
Speaker: Elise Wang (California State University, Fullerton) -
25
Insights from an Unsound Experiment: Testing Kraken, TrOCR, and VLMs with an LLM Judge
As Automatic Text Recognition (ATR) pipelines rapidly evolve, the digital humanities are confronted with a highly fragmented landscape of transcription technologies. Researchers today must navigate between specialised, fine-tuned line-level engines, such as Kraken, TrOCR, and the emerging zero-shot capabilities of generalised Vision-Language Models (VLMs). All approaches have different advantages and disadvantages which can be leveraged. However, as generating transcriptions across multiple engines becomes easier, evaluating their true quality on complex historical documents remains a significant hurdle. Traditional string-matching metrics like Character Error Rate (CER) are overly rigid, often penalising minor orthographic variations while failing to assess whether the core semantic and structural meaning of the text was successfully conveyed.
To test the limits of modern evaluation, this talk presents the results of a deliberately "unsound" experiment built on such an inference pipeline. We pitted highly specialised, fine-tuned ATR engines against unconstrained, zero-shot VLMs and deployed a Large Language Model (LLM) as an automated judge to qualitatively assess the resulting transcriptions.
Methodologically, this setup is inherently flawed: it compares apples to oranges in terms of model architecture, while relying on an LLM judge that is intrinsically biased toward modernising historical spellings and hallucinating missing context. Yet, despite these structural flaws, the experiment yielded critical insights. We will discuss the practical realities of deploying multi-engine ATR inference, expose the distinct error patterns unique to VLMs versus „traditional“ models, and demonstrate what the failures of an „LLM as a judge" reveal about the future of benchmarking historical text recognition.
Speaker: Tobias Hodel (University of Bern)
-
22
-
Working Group 3: Transcription approaches II (Transcription decisions, their effect, and related terminology) Room 2
Room 2
-
26
Otto Dei gram, rogat vram Clam. From facsimile and descriptive transcriptions to “quasi-diplomatic” and interpretative approaches in ATR/HTR
The rise of automated/handwritten text recognition has revived an old discussion that had seemed otherwise dormant: How should historical texts be transcribed? The main tension is between two approaches: the first calls for true-to-character transcriptions, while the second purposefully modifies the text, following long-established philological traditions. There are many confusing names used for these approaches: facsimile transcription, transliteration, graphemic transcription, variously prefixed diplomatic transcriptions (hyperdiplomatic, semidiplomatic), as well as interpretative or normalized transcriptions. The first part of the paper will attempt to provide an overview of them.
The main focus of the presentation, however, will be on the essence of these approaches and the justification they use in dealing with the most contentious issue: the expansion of Latin abbreviations. The arguments for the true-to-character approach focus on the effectiveness of AI training and a clear division of tasks, where the expansion is left for further processing. This promises transparency of the results. On the other hand, philological traditions have purposely moved away from true-to-character transcriptions in their evolution, preferring to devote expertise to expanding abbreviations in order to create readable texts.
I will leave aside the technological advantages of both approaches—noting only that both can achieve good results—and I will focus on methodological issues. I will contextualize both approaches not only from the current perspective but also from historical ones, and I will question whether they should really be seen as two opposites: one supposedly transparent and devoid of interpretation, and the other interpretative.
Speaker: Jan Odstrčilík (Institute for Medieval Research, Austrian Academy of Sciences) -
27
Lost in Transcription? Reconciling HTR Output with Old Czech Editorial Standards
Handwritten text recognition (HTR) is usually judged by how faithfully it reproduces what stands on the page. Czech editorial practice asks for something else. Old Czech orthography leaves consonants and vowel quantity underdetermined, so a single written form regularly supports several grammatically sound readings. Vowel quantity in Czech carries grammar, not just sound. An edition or a linguistic corpus is expected to resolve that: to commit to one reading and to say so. HTR output reproduces whichever convention its ground truth encoded, without registering that a convention was applied at all. The result is a gap that accuracy figures do not measure: a model can score well and still produce text that no Czech editorial tradition would accept.
This talk asks how far that gap can be closed, and at what cost. The models currently available for Old Czech are diplomatic almost without exception, which leaves two routes: train on interpretative ground truth and accept that the interpretation becomes invisible inside the model, or keep the output diplomatic and rebuild the editorial layer afterwards, where the decisions stay explicit and reviewable. Neither is free. I set out what each route makes visible and what it hides, and argue that the question editors should be asking is not how accurate a model is, but which decisions it has already made on their behalf.
Speaker: Anna Michalcová (Czech Language Institute, Czech Academy of Sciences; Institute for Medieval Research, Austrian Academy of Sciences) -
28
Reading Glagolitic in the Twenty-First Century: Between Philology, HTR, and AI
The digitization and scholarly editing of Glagolitic texts are entering a new phase in which traditional philological practices increasingly intersect with automated text recognition and artificial intelligence. This paper reflects on recent experiences with the transcription, transliteration, and digital editing of Croatian Glagolitic manuscripts, focusing on the opportunities and limitations of current HTR and AI-based tools.
Existing HTR models have proved useful for certain types of Glagolitic, particularly the formal angular Glagolitic found in many liturgical manuscripts. Their applicability to non-liturgical manuscripts is considerably more limited, owing to the diversity of letterforms and writing styles, while reliable models for cursive Glagolitic are still largely lacking. Recent experiments with large language models likewise show that, despite their success with other historical scripts, they generally struggle to recognize Glagolitic.
These technological limitations intersect with a fundamental philological problem: Glagolitic texts are characterized by considerable linguistic and orthographic variation and do not conform to a single standardized transcriptional tradition. Models trained on particular philological practices may therefore encode assumptions that are appropriate for some texts but problematic when applied to others.
Drawing on examples from recent projects involving the digital edition of Glagolitic and Old Czech texts, this presentation considers the relationship between established philological methods and emerging technologies. The main question is: should traditional philological practices be adapted to the capabilities of available technologies, or should digital tools instead be developed to accommodate the requirements of philological practice? Ultimately, the paper argues for a dialogue between the two, in which technological solutions remain responsive to the linguistic, textual, and editorial diversity of historical sources.
Speaker: Ana Mihaljević (Institute for the Croatian language) -
29
HTR of Normalized Latin Texts: Insights from Liturgical Manuscripts
Nearly eleven hundred sacramentaries copied up to c. 1100 — complete codices, fragments, and palimpsests — survive, dispersed across collections worldwide. The scale of this corpus has kept it outside granular analysis, and the tradition is still described through a handful of types established by earlier scholarship. RITUS+, a tool under development, addresses that scale by treating transcription not as an end but as a query.
The pipeline runs from images or IIIF manifests through Kraken-based recognition to automatic liturgical indexing. Colour analysis separates red from black within each line, so that rubrics are recovered as structure rather than merely as text. Recognition output is then tokenized and matched — by iterative fuzzy search with progressively relaxed thresholds, refined by exact Levenshtein distance — against a curated lexicon of prayers. Each identified formula receives an identifier pointing to the standardized text in the database, not to the manuscript's own spelling.
This normalization is deliberate, and it carries a cost: the tool does not address philological questions about individual variants. What it buys is that an imperfect transcription remains a serviceable query, since for a genre this formulaic the operative unit of recognition is the formula rather than the character. Indexing then permits comparison on content and on sequence, filtering a manuscript against the better-known indexed traditions and isolating what belongs to none of them — the material still awaiting scholarly attention.
Speaker: Paweł Figurski (Polish Academy of Sciences)
-
26
-
5:00 PM
Break
-
Keynote Room 1
Room 1
-
30
New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and Interpretation
Generative AI has proved itself on many tasks where we can easily evaluate its performance, from transcribing modern documents and speech to linguistic annotation to mathematical proofs. To make progress on deeper research tasks in the humanities, new benchmarks and new methods of model interpretability should be developed. We present our work on the Apollo project for developing large models for textual and historical scholarship. Compiling the largest open-access dataset for ancient Greek (with Latin in progress), we train decoder models to restore texts with long, variable, and unknown gaps and surpass the performance of existing approaches that require exact estimates of lacuna size. We evaluate models on both agreement with published restorations and scholarly feedback on unrestored texts. Our models for Greek establish the foundations for scalable archival transcription of Greek historical documents, a key research objective of the Apollo project. We describe our approach to transcription and to linking transcription and restoration to interpretable evidence. Going beyond prediction to interpretation is necessary to advance our goals in philology, the study of historical documents, and the development of more broadly applicable AI.
Zoom link for the keynote: https://oeaw-ac-at.zoom.us/j/66241437371?pwd=TCk0D1DcZoZ5DcTORa9lPsCcFe6jbi.1
Speakers: Anna Dolganov (Austrian Archaeological Institute, Austrian Academy of Sciences), David Smith (Northeastern University)
-
30
-
6:30 PM
Reception
-
8:00 AM
-
-
Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (talks) Room 3
Room 3
-
31
Testing the Reuse Value of Pracalit Ground Truth for Bhujimol Manuscripts
This presentation examines the applicability of ground truth for handwritten text recognition (HTR) on Nepalese manuscripts written in the Pracalit script (which is commonly used for Sanskrit and Newar manuscripts dating from approximately the 16th to 20th centuries) to use on an earlier Nepalese script, Bhujimol (which was prevalent from the 11th to 17th centuries). The presentation will outline several strategies identified as effective for adapting ground-truth data from one script to a similar script, and will discuss the broader relevance and significance of such models for scholarly research and cultural preservation.
Speaker: Alexander O'Neill (Musashino University) -
32
From Old Icelandic HTR Outputs to Normalised Texts through Seq2Seq Transformers
For an Old Icelandic manuscript-to-edition pipeline, we train a 10M-parameter PyTorch CharSeq2Seq Transformer to transform facsimile-like transcriptions into diplomatic transcriptions (facs2dipl task) and, in turn, into normalised ones (dipl2norm task; https://huggingface.co/NKCZ/old-icelandic-facs2dipl2norm).
We train the model on around 30,000 line-level triples of facsimile-like, diplomatic, and normalised transcriptions from three MENOTA editions of manuscripts edited by Andrea de Leeuw van Weenen. The facsimile-like transcriptions from these editions also constitute a substantial part of the training data for the OICEN-HTR model.
We evaluate CER on in-domain test set and on 200 lines from two out-of-domain manuscripts, comparing our model with frequency-based lookup tables and GPT-5.6 Luna in zero- and few-shot settings. We separately fine-tune the model on 500 lines from each out-of-domain manuscript. The base model outperforms the alternative methods on in-domain test set(CER 0.01 and 0.03), while the fine-tuned models perform best on their respective out-of-domain manuscripts (CER 0.05 and 0.12; 0.08 and 0.07). Finally, applied to HTR output from the Kringla fragment (HTR CER 0.16), the model achieves CERs of 0.20 and 0.22 for the two transformation tasks."
Speaker: Nikola Krisztian Czindrity (University of Vienna) -
33
Using OCR to expand a treebank of historical Yiddish: A Progress Report
The Penn Parsed Corpus of Historical Yiddish (PPCHY) is a treebank — text annotated with part-of-speech and syntactic information — developed for studying syntactic change in Yiddish. While PPCHY has been a valuable resource, at roughly 200K words it is quite small. Recent work has begun expanding the treebank, both with more modern (20th-century) text and with older material printed in Vaybertaytsh, a semi-cursive script used from the 16th to early 19th century. Substantial additional material survives in this script, and one goal of this work is to automatically annotate OCR'd texts, vastly increasing the treebank's size. This talk gives an overview of that work, including initial experiments applying OCR systems to Vaybertaytsh source texts in the treebank.
Speaker: Seth Kulick (Linguistic Data Consortium, University of Pennsylvania)
-
31
-
Working Group 1: Roundtable Room 1
Room 1
-
Working Group 2: Handwriting Classification I Room 2
Room 2
-
34
From Handwriting Classification Errors to Paleographic Evidence: Graphic Compensation in Ancient Greek Documentary Hands
Handwriting classification is usually evaluated through predictive accuracy, yet the structure of its errors can also provide evidence about the visual organization of historical scripts. In this work, we analyze a ConvNeXt-V2 classifier for handwritten ancient Greek letterforms trained with lacuna-based fragmentation and dynamically learned supervised contrastive loss. Rather than treating misclassifications as noise, we study recurring pairwise confusions as indicators of visual proximity between letterforms. Applied to Hellenistic Greek papyri, this analysis reveals strong confusions between visually similar but phonetically distinct letters, including Alpha–Lambda, Omicron–Theta, Iota–Rho, and Tau–Upsilon. These patterns are consistent with Irigoin’s hypothesis of graphic compensation, according to which graphic systems tend to preserve visual distinctiveness between letters with similar phonetic functions. The tendency remains stable across alternative training objectives, papyrus-level splits, and damaged-character evaluation. More broadly, the study shows how handwriting classification can be used not only for recognition, but also as a quantitative analysis of paleographic structure.
Speakers: Asimina Paparrigopoulou (Democritus University of Thrace), Paraskevi Platanou (Athens University of Economics and Business, and Archimedes, Athena Research Center, Greece) -
35
Classification of Armenian manuscripts without OCR/HTR preprocessing
Armenian manuscripts are a particularly rich corpus when it comes to their premodern “metadata” - the addition of a colophon, usually naming the scribe and often giving the date as well as circumstantial information about the production of the manuscript, was a strong element of the copying tradition. Moreover, when it exists, this metadata is usually reproduced in the manuscript catalogue entries. While there do exist models for HTR transcription of classical Armenian, in this project we are interested in what can be modelled without the time-consuming process of HTR analysis and correction. We attempt to identify approximate dates of copying for manuscripts that are missing a colophon, and use the results of this analysis to revisit the timeline of Armenian script development.
Speaker: Tara Andrews (University of Vienna) -
36
Beyond the Heatmap: What Faithful Explanations Can Do for Handwriting Identification
Explainable AI for handwriting identification has so far mostly stopped at the heatmap: pixel-level relevance maps that scholars can inspect but rarely act upon. This talk traces a path from diagnosis to action to discovery. We first present a validated, model-agnostic explainability framework for medieval handwriting classifiers, assessed for faithfulness, stability, and cross-model agreement, and benchmarked against expert paleographic judgement across four Latin manuscript datasets. We then show that these faithful explanations can be repurposed as data-curation signals: a data-centric retraining strategy that uses positively-attributed evidence to construct improved second-stage training sets, systematically outperforming random sample selection on Vatican manuscript datasets. Finally, we outline how this foundation feeds into an ongoing research programme that aims to reframe handwriting identification as an explainable, open-set discovery process at the scale of the Codices Latini Antiquiores — moving from patch-level saliency to segment-based explanations, aggregated graphic signatures, and natural-language descriptions grounded in a paleographic ontology. Together, these strands argue for a shift in how explainability is used in digital paleography: not as an afterthought to classification, but as an instrument that can curate training data and, ultimately, support historical interpretation itself.
Speakers: Serena Ammirati (Università degli Studi Roma Tre), Paolo Merialdo (Roma Tre University)
-
34
-
10:30 AM
Coffee Break
-
Working Group 4: Roundtable 🎤 - Language Challenges Room 1
Room 1
-
Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (demos) Room 3
Room 3
-
37
New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer
The bibliographer and literary critic Jerome McGann has written that “original documents are fictions we practice in order to manage their losses and our limits” (A New Republic of Letters, 5). This talk considers how a corpus of early twentieth-century Korean novels and periodicals is being practiced as a new kind of fiction within prototype software called Mo文oNExplorer. It describes how a loosely organized team comprising an American Koreanist, a Korean small business owner, a Korean startup founder, and a Vietnamese programmer attempted to engineer contexts of use for older Korean texts: presenting them again as digital transcriptions, as translations into contemporary Korean for present-day readers, and as indexed passages resituated and summarized through a RAG system. The aim was to create a fiction that might someday be practiced usefully within institutional settings such as the National Library of Korea.
Speaker: Wayne de Fremery (Dominican University of California) -
38
From Document Images to Research Catalogue
This demonstration builds on our experience with 19th-century archival documents from Chocó, Colombia. As researchers digitised these documents, there was a concurrent need for datafication to assess the collection's scope, content, and research potential. We developed a minimal tool to extract text using vision-language models (VLMs), identify and normalise entities, and create a search interface. This catalogue tool, named ficherito, addresses a common use case in which researchers need machine-readable text and structured data for exploratory data analysis. This pilot identified the need for more full-featured software, called fichero, demonstrated by Daniel Tubb.
Speaker: Andrew Janco (Princeton University) -
39
Fichero and the Circuit Court Archive of Istmina: An AI Workflow App for Researchers
In this talk, I demo Fichero. Fichero is a macOS app, a SwiftUI front end over a FastAPI Python engine, that puts machine-learning and AI workflows, using local and hosted models, in researchers' hands. Developed from a collaboration with Andy Janco (Princeton) and Ann Farnsworth-Alvear (UPenn), I demonstrate Fichero on the Circuit Court Archive of Istmina, Colombia. The archive is a British Library Endangered Archives Programme project consisting of 61,000 documents of Afro-Colombian legal records (1870–1930) from the Chocó, Colombia. These handwritten and typed documents are often faded and damaged, divided into 834 cases, and contain material on mining, land disputes, workplace injustices, and much more. Fichero allows researchers to import the archives as folders and images, use a Mac interface to enhance degraded images (deskew, denoise, and adjust contrast with OpenCV); transcribe handwritten and typed pages (OCR/HTR with Apple Vision, open-source, and frontier vision-language models); run (and visually edit) AI workflows powered by LangGraph and LangChain to break steps down to produce metadata; extract people, places, and subject–verb–object claims into a knowledge graph (named-entity and relation extraction via Apple Intelligence, spaCy, or local and hosted LLMs); write catalogue and finding-aid entries (using frontier models); search by meaning (vector embeddings in LanceDB alongside a DuckDB graph index); query the corpus through GraphRAG chat (retrieval over the graph plus vectors, with cited sources); compare model outputs; and publish a catalogue (an 11ty site with search). Fichero offers an easier to use tool so that skilled archivists, researchers, and others can work with archives, deploy workflows, and use AI and machine-learning tools and techniques.
Speaker: Daniel Tubb (Anthropology, University of New Brunswick, Canada) -
40
Intertextuality as a Retrieval Task: Benchmarking Text Reuse in Classical and Medieval Latin
What if we reframe intertextuality detection as a retrieval task? In this talk, I will take quotations from an author and rank them against corpus of possible sources. This allows me to measure recall@10, precision@1, nDCG@10 and compare four different methods that are commonly used: BM25, dense embeddings (generic and Latin-specific), reciprocal-rank fusion, and reranking by cross-encoder or large LLM. Due to the lack of a gold standard for medieval Latin, I used classical Latin (Loci Similes; Jerome and Lactantius against ~90,000 passages), where verified links are available as a published benchmark, and created my own for medieval Latin (Bernard of Clairvaux against the Vulgate), where no such benchmark exists and the reference set has to be assembled from a machine-parsed Patrologia Latina apparatus and a BiblIndex export. Results show that no single method wins across corpora. Fusion of BM25 and an embedding model has potential to give the best coverage on both corpora and reranking lifts an already strong lists. Part of this benchmarking experiment was a discovery run on Bernard. The pipeline proposed 934 candidates, of which 308 were known references independently re-found; of the remaining 626, a machine verifier flagged 166 for reading, and hand review confirmed 28 previously unrecorded scriptural links — a third of them invisible to word-matching, and several of them explicitly introduced, word-for-word quotations recorded in neither Migne nor BiblIndex. Such systems work best as calibrated assistants whose "false positives" are often worth following.
Speaker: Martin Roček (Institute for Medieval Research, Austrian Academy of Sciences and Faculty of Arts, Charles University)
-
37
-
Working Group 2: Handwriting Classification II Room 2
Room 2
-
41
Classifying Squeezes Again: Initial Results from an ICDAR Contest
At the first SCOOP meeting in June 2025 Aaron Hershkowitz and Nicholas Howe presented the results of their 2024-25 NEH-funded project to evaluate various approaches to automated text recognition. Two main text recognition approaches were explored: traditional sequential Optical Character Recognition (OCR), and Scene Text Detection. Out of the box OCR engines like Kraken were found not to handle squeeze images well, even after processing to convert them to “pseudo-ink”. Sequential OCR performed much better when retrained on squeeze images, but still lagged behind the performance of Scene Text Detection models like YOLO and DBNet. The most successful approach turned out to be using Scene Text Detectors to find characters, then assembling lines from those characters for sequential approaches like Convolutional Neural Networks. In 2025-26 Hershkowitz and Howe ran a contest through the International Conference in Document Analysis and Recognition (ICDAR) to attract outside researchers with fresh approaches to the problem. That contest concluded at the end of April, and initial analysis of the results has been conducted. As part of WG2, Aaron Hershkowitz will present the team’s conclusions from the ICDAR contest, and discuss potential next steps in terms of improving and utilizing squeeze ATR.
Speaker: Aaron Hershkowitz (The Institute for Advanced Study) -
42
Communities of Practice: Capturing the Aspect of Late Medieval Handwriting
Communities of Practice is funded by the Social Sciences and Humanities Research Council of Canada and explores how the multilingual urban bureaucracy of London — the world of Chaucer and his contemporaries — shaped late medieval literature. Bringing together linguistic, palaeographical, literary-historical, and computational methods, the project examines how networks of clerks employed in government offices contributed to the production, circulation, and transformation of literary texts between 1377 and 1471. By pairing innovative machine-learning techniques with traditional manuscript study, and making full use of the increase in digital access to archival documents, the project aims to move beyond linking single documents with individual clerks. Instead, it works towards a broader understanding of the distinctive scribal practices that arose in particular office communities and the role played by these scribal communities in the development of literary culture.
Scribal handwriting identification poses an unusual machine-learning problem: ground truth is largely unrecoverable. Few medieval scribes signed their work, and while palaeographers have attributed many documents with broad consensus, the contested cases, including some of the most consequential for literary history, resist resolution. Modern writer-identification methods achieve near-perfect accuracy on contemporary handwriting but degrade sharply on historical material. Meanwhile, digitisation has produced vast quantities of unlabelled manuscript images. In Communities of Practice we explore self-supervised representation learning to exploit this unlabelled data, with the goal of building tools that capture the “aspect” of writing, rather than opaque scores or classifications. The aim is not to settle attributions but to give scholars a new instrument with which to argue them.
Speakers: Sebastian Sobecki (University of Toronto), Sam Grieggs (Indiana University of Pennsylvania) -
43
Script classification and alphabet identification
Handwritten Text Recognition assumes what encrypted historical manuscripts withhold: a known script, and labelled examples of it. Two questions therefore come first. Which script family does this document belong to, and which symbols does it actually use? Both are still answered by hand, an expert-bound process that does not scale to collections of hundreds of manuscripts. This talk reports what happens when standard computer vision is turned on those questions, and why it fails in an instructive way. An unsupervised pipeline that segments, embeds and matches symbols against a database of reference alphabets assigns a cipher to the wrong alphabet with 76% confidence, even though its true script is absent from that database altogether. The winning reference is simply the only one cropped from a real manuscript instead of rendered from a font. Style, not shape, dominates the embedding space. I trace the consequences of that diagnosis: adversarial domain alignment and inference-time style adaptation, which yield a label-free script fingerprint consistent with palaeographic knowledge; the same entanglement along a temporal axis, across a millennium of Greek letterforms; and style deliberately reused as signal, for scribal attribution and dating.
Speaker: Giuseppe De Gregorio (Computer Vision Center - CVC - Barcelona)
-
41
-
12:30 PM
Lunch
-
Working Group 5: Datasets and Institutions Room 2
Room 2
-
44
Integrating HTR into Digital Library Workflows
With the rapid improvement of handwritten text recognition models, large language models, and multimodal AI models, the University of Pennsylvania Libraries is assessing how these technologies can be applied at scale across its handwritten collections. This presentation will focus on an emerging project to apply HTR to a group of fifteenth-century Italian manuscripts written in humanistic scripts. It will explore challenges and opportunities relevant to the entirety of the Libraries’ multilingual manuscript collection and consider how transcription datasets resulting from HTR projects can be leveraged for different library applications.
Speaker: Jessie Dummer (University of Pennsylvania Libraries) -
45
Manuscriptorium Full-Text Module: Building a Digital-Edition Infrastructure
After several years of development, the transition to the new version of the Manuscriptorium digital library was completed last autumn. The development, however, is still ongoing – the most significant and recent part of which is the TEI-compatible full-text module. As a result, Manuscriptorium offers its users an alternative core that shifts the central perspective from visual representation of digitised documents to textual content, with images serving as a complement to the texts they contain.
The overall objective is to build a solid and stable infrastructure, enabling the community of scholars as well as the general public to read, study, and research the historical texts within a wider Central European digital context. At present, the main focus lies on converting older editions of historical texts of Bohemian origin. The editions are then correlated not only with the documents forming their textual basis but also with other known witnesses, thereby reviving such editions in a broader and more dynamic context. In the extended process of establishing an exact method of data preparation, their subsequent processing, and visualisation workflows, we have come across several methodological problems, primarily in relation to directly connecting digitised fragmentary witnesses of varying extent and content with the full texts of the editions.
From the end-user perspective, the edition module consists of two components: an application for advanced viewing of the texts integrated with the digitised images, and a catalog enabling high-granularity searching across the edition database. The essential data persistence has already been secured by the overarching Manuscriptorium resolver infrastructure. By means of full IIIF compatibility, the texts are correlated with the digitised documents from Manuscriptorium's content core as well as from external image repositories. Parallel yet differentiated browsing of texts and images, along with the additional visualisation of modern translations has also been successfully implemented. The next developmental phase will be primarily dedicated to processing and visualising complex critical apparatuses. On the whole, compared with other historical-text databases, our principal purpose is to provide users with highly curated – and thus reliable – textual content and to make it as accessible as possible, not least through an engaging and easy-to-navigate user interface.
In addition, we are taking the first steps towards the inclusion of HTR-generated transcriptions in the database and, ultimately, in the wider Manuscriptorium environment. One goal is for these transcriptions to serve as a complementary dataset, allowing users, for instance, to visualise raw HTR data representing individual witnesses alongside the textual constructs resulting from critical editorial work. In the case of revised and moderately enriched HTR transcriptions, we also aim to incorporate them directly at the level of pragmatic editions, thereby further diversifying the edition module’s content.
Speaker: Michael Lužný (National Library of the Czech Republic) -
46
Reviving the Legacy: How to Adapt the Latin Text Archive for the Age of AI
The Latin Text Archive (LTA) is an online platform hosted by the Berlin-Brandenburg Academy of Sciences (BBAW) since 2020 (https://LTA.bbaw.de). Its primary objective is to facilitate computer-assisted semantic analysis of Latin texts and corpora spanning various epochs and genres. Conceived as an open platform from its inception in the mid-2000s, it enables external editors and text providers to make their editions available for morphological and semantic analysis. This is why it currently contains over 58,000 texts comprising more than 106 million words. However, the LTA is currently facing three significant challenges: 1) How can metadata from different editions and academic backgrounds be aligned? 2) How to organise the text material for AI-supported text analysis. 3) How can the LTA be maintained and continuously improved in times of reduced funding? This presentation addresses these three challenges and aims to stimulate discussion by exploring possibilities rather than presenting definitive solutions.
Speaker: Tim Geelhaar (Goethe Universität Frankfurt am Main) -
47
HTR in the German Manuscript Centres: Current Status, Challenges, and Perspectives
The German Manuscript Centres have a long-standing tradition of developing standards, infrastructures, and services for the scholarly description and accessibility of medieval manuscripts. With the transition from printed catalogues to digital infrastructures such as Manuscripta Mediaevalia and, since 2023, the Handschriftenportal, digitisation has fundamentally expanded both the possibilities and responsibilities of the centres.
Against this background, the German Manuscript Centres are increasingly engaging with HTR. Their current activities focus on three main areas: establishing standards and guidelines for the creation and documentation of Ground Truth data; providing high-quality, openly accessible datasets and models for medieval scripts; and developing workflows for integrating manuscript full texts into the cross-institutional infrastructure of the Handschriftenportal and making them searchable.
Particular attention is being paid to interoperability, metadata, quality assurance, and different levels of transcription.HTR thus represents not a replacement for traditional manuscript expertise, but an important extension of it. The Manuscript Centres increasingly see themselves as centres of expertise for the digital documentation and study of historical manuscripts, combining technological innovation with palaeographical and codicological expertise to ensure the quality, sustainability, and scholarly usability of digital resources.
At the same time, close collaboration with the research community remains essential. In particular, the collaborative creation, correction, and enrichment of full-text transcriptions can help bridge the gap between technological infrastructure and scholarly research. Such collaboration ensures that HTR workflows are not only technically robust, but also responsive to the needs, practices, and research questions of scholars.
Speaker: Ursula Stampfer (Bayerische Staatsbibliothek)
-
44
-
Working Group 6: Unconference Room 3
Room 3
-
Working Group 3: Building ATR/HTR pipelines Room 1
Room 1
-
48
The Pragmatics of ATR Pipelines: Structural Dilemmas in Designing Workflows for Historical Corpora
Automatic Text Recognition has become the most widely adopted application of machine learning in the humanities, yet projects often encounter challenges when integrating existing models into their workflow. This paper dicusses the source of these challenges as a mismatch between requirement profiles - a concept broader than transcription guidelines, constituted by three dimensions: disciplinary standards, downstream implementation requirements , and the „Erschließungsmodus“, the degree of acceptable intervention by a model. Since ATR models are data-deterministic, current Research hast to choose between two strategies: heterogeneous „super models“ offering coverage but also unpredictable output, or convention-consistent models that are predictable but rigid, costly to produce, and move the problem into downstream postprocessing.
The paper proposes a third option: encoding the requirement profile in the model itself. Extending the language-token mechanism of Benjamin Kiessling's party architecture with two conditioning tokens (corpus: transcription convention and character inventory, and mode: abbreviated / expanded), and replicating the approach in a CNN based model, experiments on data from Burchards Dekret Digital show that transcription behaviour can be steered reliably: expected abbreviation markers appear in 83–86% of cases in abbreviated mode and are suppressed in 99–100% of cases in expanded mode, with steerability persisting when the training pool is enlarged by CATMuS and Tridis material and when BDD's token is forced onto unseen corpora.
Building on this proof of concept, the paper argues for a shared, machine-readable taxonomy of requirement profiles that would standardise the description of transcription practice rather than the practice itself.
Speaker: Michael Schonhardt (TU Darmstadt / Akademie der Wissenschaften und der Literatur | Mainz) -
49
Building an ATR pipeline for tabular data and script translation
My presentation addresses the technical challenges of applying automated text recognition (ATR) to demographic data preserved in late Ottoman census registers. It builds on the experience we gained in the LOOP research project—Late Ottoman Palestinians. Two issues are central. First, the registers are handwritten in Ottoman Turkish, an under-resourced historical language for which effective pre-trained ATR models remain scarce. Second, the data consists of handwritten entries embedded in tabular forms, many with empty cells and others containing multi-line content that challenges standard layout assumptions. Existing table-segmentation methods perform poorly on such material, resulting in unreliable region detection and reduced transcription accuracy.
To address these challenges, we rethought segmentation entirely. Instead of attempting table recognition, we trained the segmentation model to draw a single line spanning the full width of each table row. During text recognition, vertical column boundaries are allowed to be recognized as pipe characters. In post-processing, the table is reconstructed from the text lines with the pipe character as a separator, as if from a CSV file. This approach achieved an overall text recognition rate of 94%, with an error rate of under 2% for years of birth. A further challenge emerged in segmentation: while line placement can be trained, the associated masks that define the text recognition area cannot. Due to empty cells and multi-line entries, these masks are often positioned too low. To address this, we export the segmented files and apply a custom Python script that generates new masks aligned to the full height of each table row. The corrected files are then re-imported for text recognition.
I speak from the perspective of end users of digital humanities tools. Rather than proposing new algorithms, we combine established platforms and models and extend them through Python scripts to construct a robust processing pipeline. Since few components functioned “out-of-the-box” for this material, the process required iterative experimentation and, at several points, deliberately “outside-the-box” technical decisions.
Speaker: Olaf Berg (Ruhr-Universität Bochum) -
50
Iterative HTR: A new pipeline for automatic text recognition of Arabic-script
We study word-level text-to-image mapping in Arabic-script manuscripts: given a line image and its transcription, locate each word. At corpus scale, accurate word geometry opens new possibilities for digital paleographical analysis while improving ATR both upstream and downstream. This mapping is learned from 25,000 line crops drawn from 87 manuscripts, supervising word geometry only through the transcription, and evaluated against a human-drawn gold-standard set of 1,276 words. We compare against a blind control that never sees the transcription.
Speaker: Osama Eshera (University of Maryland) -
51
Building an OCR Pipeline at the Austrian National Library
One of the Austrian National Library’s Strategic Goals for 2023–2027 is to create new ways for users to explore its collections. The improvement of ATR, especially OCR, falls within this context. During my presentation, I will talk about the currently ongoing OCR Project (2024–2027) and one of its main outcomes, an internal OCR Service. This Service contains a modular OCR pipeline using open-source models, that was built by ONB’s software developers. This pipeline has already produced OCR for a major 3rd party funded digitization project, “Viele Stimmen,” which entailed recognizing text on historical pamphlets with challenging layout and font. Finally, I will explain how the pipeline and the Service will be used in the future and discuss past and future challenges.
Speaker: Johannes Knüchel (Austrian National Library)
-
48
-
3:00 PM
Coffee Break
-
52
Working Groups working on conclusions
- WG 1 [Room 1 = HS1]: HTR Technology Development (coordinators: Tobias Hodel)
- WG 2 [Room 2 = SR5]: Document or Handwriting Classification (coordinators: Maria Konstantinidou, John Pavlopoulos, Paraskevi Platanou)
- WG 3 [Room 3 = SR4]: Methodological Issues of HTR (coordinators: Anna Michalcová and Jan Odstrčilík)
- WG 4 [Room 4 = SR3]: Language Challenges (coordinators: Christine Roughan)
- WG 5 [Room 5 = HS2, second floor]: Datasets and Institutions (coordinators: Gerda Heydemann,
Maria Konstantinidou) - WG 6 [Room 6 = HS3, third floor]: Leveraging Outputs: Text Reuse, NLP, and More (coordinators: Tara Andrews, Martin Roček)
-
5:00 PM
Short Break
-
53
Final Roundtable - Open data, open code, open minds in AI
-
54
Conclusion of the public part
-
-
-
55
Internal meeting of SCOOP
-
55