Contribution List

55 out of 55 displayed
Export to PDF
  1. Thibault Clérice (ALMAnaCH, Inria)
    9/7/26, 9:15 AM

    Automatic Text Recognition (ATR) has made manuscript scholarship scalable, but it also intro- duces profound epistemic challenges stemming from the conflation of transcription (the reporting of the visible source text) and edition (the act of scholarly interpretation). We argue that some of the current ATR practices may produce silent forgeries—plausible yet undocumented inferences that...

    Go to contribution page
  2. William Mattingly (Yale University)
    9/7/26, 11:00 AM

    The rapid evolution of frontier Vision-Language Models (VLMs), such as the Gemini 3.5 family, has introduced unprecedented zero-shot capabilities in complex multimodal tasks, including high-fidelity document transcription and precise spatial reasoning via bounding box generation. As these massive, proprietary models increasingly solve general visual-linguistic tasks without task-specific...

    Go to contribution page
  3. Colin Brisson (Ecolé pratique des hautes études)
    9/7/26, 11:20 AM

    This presentation introduces AnandaSky, a compact vision-language model developed for line-level transcription of historical sinographic documents, and reports ongoing work toward a new page-level version designed for broader cross-domain use. AnandaSky combines a shallow, high-resolution visual encoder with a Qwen3-0.6B autoregressive decoder. Global visual attention, 10-pixel patches, an...

    Go to contribution page
  4. Andy Stauder (READ COOP)
    9/7/26, 11:40 AM

    It seems the age of thinking in tools is at an end. While Large Language Models offer rapid text processing, generic chat interfaces struggle with complex layout analysis, historical scripts, and verifiable ground truth, and especially control over workflows. As a toolbox, Transkribus follows a different path, aiming for a necessary alternative and an essential complement to LLM-driven...

    Go to contribution page
  5. Benjamin Kiessling (ALMAnaCH, Inria Paris)
    9/7/26, 1:30 PM

    Computational methods have transformed many areas of historical research, yet
    paleography has seen comparatively few algorithms and tools that can operate at
    scale while remaining adaptable to scholarly practices across different
    scripts. This is not due to a lack of interest in digital methods. The growth
    of digital repositories, annotation platforms such as DigiPal, and...

    Go to contribution page
  6. Ana Mihaljević (Institute for the Croatian language)
    9/7/26, 1:30 PM

    The digitization of the Dictionary of the Croatian redaction of Church Slavonic, carried out within the DigiSTIN project, has presented us with a particular challenge: how to efficiently OCR and HTR materials that combine multiple languages and writing systems. The dictionary covers Croatian Church Slavonic, Croatian, English, Latin, and Greek, and uses Latin, Greek, Glagolitic, and Old...

    Go to contribution page
  7. Doug Emery (University of Pennsylvania Libraries)
    9/7/26, 1:30 PM

    Since 2023, members of Penn Libraries’ Cultural Heritage Computing, Schoenberg Institute for Manuscript Studies, and Research Data and Digital Scholarship groups have conducted workshops and pilot projects to explore workflows for handwritten text recognition (HTR) and its potential to support transcription of library collections for discovery, access, and computational use. This talk will...

    Go to contribution page
  8. Bernhard Bauer (University of Graz)
    9/7/26, 1:40 PM

    The early medieval period in Western Europe was marked by intense linguistic and cultural interaction, as vividly attested in the glossed manuscripts of the era. These glosses are found in the margins or between the lines of Latin texts. They provide unparalleled evidence of the multilingual environment in which manuscripts were produced, copied, and read. The ERC-project GlossIT addresses the...

    Go to contribution page
  9. Baptiste Queuche (Calfa)
    9/7/26, 1:50 PM

    Calfa is developing general and custom AI models for researchers and cultural heritage professionals, dedicated to recognize the text and analyze documents in non-western languages, such as Arabic, Armenian, Georgian, Syriac, Chinese, Greek. Calfa also offers many open-source datasets for Arabic or Armenian languages.

    We recently developed a promising hybrid approach, combining our specific...

    Go to contribution page
  10. Wolfgang Göderle (University of Graz | University of Passau | MPI GEA), David Fleischhacker (University of Graz)
    9/7/26, 1:50 PM

    The transcription of historical administrative sources presents a distinctive set of challenges for automated text recognition: dense, formulaic layouts, highly abbreviated Latin and German terminology, and the cumulative degradation typical of large serial document corpora. This paper presents a high-precision OCR pipeline developed within the Unlocking the Schematismus project, designed to...

    Go to contribution page
  11. Sajjad Nikfahm-Khubravan (Roshan Institute for Persian Studies, University of Maryland)
    9/7/26, 1:50 PM

    Arabic manuscript traditions present a constellation of challenges that continue to resist the assumptions underlying most contemporary Handwritten Text Recognition (HTR) systems. Unlike Latin scripts, Arabic is written right-to-left in a fully cursive hand in which most, but not all, letters are connected to their neighbors, and the graphical form of each letter changes depending on its...

    Go to contribution page
  12. Seth Kulick (Linguistic Data Consortium, University of Pennsylvania)
    9/7/26, 2:00 PM

    Historical treebanks — corpora annotated with part-of-speech and syntactic information — are essential resources for linguists studying language change, but they are labor-intensive to build. Existing historical treebanks are consequently small, typically around 1-2 million words, which limits the scope of research they can support. This talk gives an overview of work that explores whether...

    Go to contribution page
  13. Alicia González Martínez (Hamburg University)
    9/7/26, 2:10 PM

    This talk presents a scalable pipeline for converting digitised Arabic books from PDFs into structured mARkdown for integration into the OpenITI corpus. The workflow begins by preparing the PDFs and processing their opening pages to extract the bibliographic metadata required to identify works and editions. The documents are then transcribed in parallel by several complementary recognition...

    Go to contribution page
  14. Malamatenia Vlachou (IRHT/CNRS-ENPC)
    9/7/26, 2:10 PM

    The analysis of medieval handwriting lies at the heart of palaeographical research. While Automatic Text Recognition (ATR) systems are primarily developed to produce transcriptions, they also generate rich data that can be leveraged to study the scripts themselves. This presentation will explore how ATR data can be exploited to develop new computational tools for palaeographical analysis....

    Go to contribution page
  15. Andrew Janco (Princeton University)
    9/7/26, 2:10 PM

    Transformer-based language models have a fixed vocabulary of tokens used to represent words and word parts. Which languages can and cannot be represented with current model vocabularies? When native support does not exist, what are the possibilities and drawbacks of byte pair encoding? This short talk presents a tool to check how well current model vocabularies support the language(s) you’re...

    Go to contribution page
  16. Isabelle Marthot-Santaniello (University of Basel)
    9/7/26, 2:20 PM

    Greek has been used in Egypt from the time of Alexander the Great to beyond the Islamic conquest. During this millennium, Greek, like any language, evolved. It was also in contact with other languages (Demotic Egyptian, Latin, Coptic, Arabic…), a multilingual situation that encouraged the emergence of loanwords and the modification of some formulas. Besides, a wide range of text types were...

    Go to contribution page
  17. Elena Chepel (University of Vienna), Anton Repushko (Independent Researcher)
    9/7/26, 2:30 PM

    In this paper, we present a new text recognition tool, Anagnostes, which
    combines modern approaches in machine learning with optical character
    recognition (OCR). It has been trained to read Greek literary and
    documentary papyri from images, with a particular emphasis on
    recognising not only neat and well-preserved fragments but also cursive
    and damaged texts. Currently, the model achieves...

    Go to contribution page
  18. Dominique Stutzmann (IRHT-CNRS / HU Berlin)
    9/7/26, 2:30 PM

    Transcription levels, the identification of scribal hands, and workflow architectures are usually discussed in three separate conversations, one editorial, one palaeographical, one technical. I fold them into two decisions taken before any computational result exists: at what grain the written surface is read (allographetic, graphematic, normalised on the text side; glyph, patch, line, page on...

    Go to contribution page
  19. Katrín Lísa van der Linde Mikaelsdóttir (University of Iceland)
    9/7/26, 3:30 PM

    Iceland’s literary heritage includes around 1000 medieval manuscripts and charters dating from the eleventh century onward, alongside several thousands of post-medieval documents. Although most of the preserved vernacular corpus has already been digitized and catalogued, Icelandic remains a low-resource language for Handwritten Text Recognition. Publicly available models are scarce due to a...

    Go to contribution page
  20. Jan Odstrčilík (Institute for Medieval Research, Austrian Academy of Sciences)
    9/7/26, 3:30 PM

    The rise of automated/handwritten text recognition has revived an old discussion that had seemed otherwise dormant: How should historical texts be transcribed? The main tension is between two approaches: the first calls for true-to-character transcriptions, while the second purposefully modifies the text, following long-established philological traditions. There are many confusing names used...

    Go to contribution page
  21. Michal Racyn (Masaryk University)
    9/7/26, 3:30 PM

    Arkindex (a document processing platform developed by TEKLIA) represents a promising open-source alternative to Transcribus and e-Scriptorium. Based on modular approach Arkindex implements machine learning algorithms that can be applied for layout analysis, transcription, structuration, named entity recognition, and information extraction of text documents as well as object segmentation,...

    Go to contribution page
  22. Tristan Repolusk (Department of Digital Humanities, University of Graz)
    9/7/26, 3:50 PM

    Medieval manuscripts, unlike modern printed documents, are typically characterized by non-uniform script and complex page layout, both with respect to larger regions (e.g., main text regions and marginal regions) and individual text lines. Especially manuscripts featuring extensive annotations such as glosses or reading signs tend to increase the page layout complexity significantly: this...

    Go to contribution page
  23. Anna Michalcová (Czech Language Institute, Czech Academy of Sciences; Institute for Medieval Research, Austrian Academy of Sciences)
    9/7/26, 3:50 PM

    Handwritten text recognition (HTR) is usually judged by how faithfully it reproduces what stands on the page. Czech editorial practice asks for something else. Old Czech orthography leaves consonants and vowel quantity underdetermined, so a single written form regularly supports several grammatically sound readings. Vowel quantity in Czech carries grammar, not just sound. An edition or a...

    Go to contribution page
  24. Christine Roughan (Princeton University)
    9/7/26, 4:00 PM

    Medieval scripts present Handwritten Text Recognition tools with a wide variety of challenges due to highly diverse orthographic features specific to different languages and scripts. Using medieval Greek and Arabic as case studies, this talk examines the handling of features from diacritical marks (accents, breathing marks, vowels, and consonant pointing) to different categories of...

    Go to contribution page
  25. Elise Wang (California State University, Fullerton)
    9/7/26, 4:10 PM

    The English common law rests on a documentary record of staggering size and frustrating inaccessibility. At least eleven million pages of medieval legal manuscripts survive at the National Archives at Kew, the product of a documentary machine at Westminster and county courts that churned out records at a rate of tens of thousands per year. For comparison, roughly six to seven thousand literary...

    Go to contribution page
  26. Ana Mihaljević (Institute for the Croatian language)
    9/7/26, 4:10 PM

    The digitization and scholarly editing of Glagolitic texts are entering a new phase in which traditional philological practices increasingly intersect with automated text recognition and artificial intelligence. This paper reflects on recent experiences with the transcription, transliteration, and digital editing of Croatian Glagolitic manuscripts, focusing on the opportunities and limitations...

    Go to contribution page
  27. Paweł Figurski (Polish Academy of Sciences)
    9/7/26, 4:30 PM

    Nearly eleven hundred sacramentaries copied up to c. 1100 — complete codices, fragments, and palimpsests — survive, dispersed across collections worldwide. The scale of this corpus has kept it outside granular analysis, and the tradition is still described through a handful of types established by earlier scholarship. RITUS+, a tool under development, addresses that scale by treating...

    Go to contribution page
  28. Tobias Hodel (University of Bern)
    9/7/26, 4:30 PM

    As Automatic Text Recognition (ATR) pipelines rapidly evolve, the digital humanities are confronted with a highly fragmented landscape of transcription technologies. Researchers today must navigate between specialised, fine-tuned line-level engines, such as Kraken, TrOCR, and the emerging zero-shot capabilities of generalised Vision-Language Models (VLMs). All approaches have different...

    Go to contribution page
  29. Ephrem Aboud Ishac (Institute for Medieval Research, Austrian Academy of Sciences)
    9/7/26, 4:30 PM

    As Automated Text Recognition (HTR) becomes increasingly integral to Digital Humanities, under-resourced scripts such as Syriac continue to present unique methodological challenges. This presentation outlines an experimental workflow designed to bridge this gap by not only recognizing Syriac handwriting but actively "giving voice" to the texts. The proposed pipeline begins with the HTR...

    Go to contribution page
  30. Anna Dolganov (Austrian Archaeological Institute, Austrian Academy of Sciences), David Smith (Northeastern University)
    9/7/26, 5:15 PM

    Generative AI has proved itself on many tasks where we can easily evaluate its performance, from transcribing modern documents and speech to linguistic annotation to mathematical proofs. To make progress on deeper research tasks in the humanities, new benchmarks and new methods of model interpretability should be developed. We present our work on the Apollo project for developing large models...

    Go to contribution page
  31. Asimina Paparrigopoulou (Democritus University of Thrace), Paraskevi Platanou (Athens University of Economics and Business, and Archimedes, Athena Research Center, Greece)
    9/8/26, 9:00 AM

    Handwriting classification is usually evaluated through predictive accuracy, yet the structure of its errors can also provide evidence about the visual organization of historical scripts. In this work, we analyze a ConvNeXt-V2 classifier for handwritten ancient Greek letterforms trained with lacuna-based fragmentation and dynamically learned supervised contrastive loss. Rather than treating...

    Go to contribution page
  32. Alexander O'Neill (Musashino University)
    9/8/26, 9:00 AM

    This presentation examines the applicability of ground truth for handwritten text recognition (HTR) on Nepalese manuscripts written in the Pracalit script (which is commonly used for Sanskrit and Newar manuscripts dating from approximately the 16th to 20th centuries) to use on an earlier Nepalese script, Bhujimol (which was prevalent from the 11th to 17th centuries). The presentation will...

    Go to contribution page
  33. Tara Andrews (University of Vienna)
    9/8/26, 9:30 AM

    Armenian manuscripts are a particularly rich corpus when it comes to their premodern “metadata” - the addition of a colophon, usually naming the scribe and often giving the date as well as circumstantial information about the production of the manuscript, was a strong element of the copying tradition. Moreover, when it exists, this metadata is usually reproduced in the manuscript catalogue...

    Go to contribution page
  34. Nikola Krisztian Czindrity (University of Vienna)
    9/8/26, 9:30 AM

    For an Old Icelandic manuscript-to-edition pipeline, we train a 10M-parameter PyTorch CharSeq2Seq Transformer to transform facsimile-like transcriptions into diplomatic transcriptions (facs2dipl task) and, in turn, into normalised ones (dipl2norm task; https://huggingface.co/NKCZ/old-icelandic-facs2dipl2norm).

    We train the model on around 30,000 line-level triples of facsimile-like,...

    Go to contribution page
  35. Serena Ammirati (Università degli Studi Roma Tre), Paolo Merialdo (Roma Tre University)
    9/8/26, 10:00 AM

    Explainable AI for handwriting identification has so far mostly stopped at the heatmap: pixel-level relevance maps that scholars can inspect but rarely act upon. This talk traces a path from diagnosis to action to discovery. We first present a validated, model-agnostic explainability framework for medieval handwriting classifiers, assessed for faithfulness, stability, and cross-model...

    Go to contribution page
  36. Seth Kulick (Linguistic Data Consortium, University of Pennsylvania)
    9/8/26, 10:00 AM

    The Penn Parsed Corpus of Historical Yiddish (PPCHY) is a treebank — text annotated with part-of-speech and syntactic information — developed for studying syntactic change in Yiddish. While PPCHY has been a valuable resource, at roughly 200K words it is quite small. Recent work has begun expanding the treebank, both with more modern (20th-century) text and with older material printed in...

    Go to contribution page
  37. Aaron Hershkowitz (The Institute for Advanced Study)
    9/8/26, 11:00 AM

    At the first SCOOP meeting in June 2025 Aaron Hershkowitz and Nicholas Howe presented the results of their 2024-25 NEH-funded project to evaluate various approaches to automated text recognition. Two main text recognition approaches were explored: traditional sequential Optical Character Recognition (OCR), and Scene Text Detection. Out of the box OCR engines like Kraken were found not to...

    Go to contribution page
  38. Wayne de Fremery (Dominican University of California)
    9/8/26, 11:00 AM

    The bibliographer and literary critic Jerome McGann has written that “original documents are fictions we practice in order to manage their losses and our limits” (A New Republic of Letters, 5). This talk considers how a corpus of early twentieth-century Korean novels and periodicals is being practiced as a new kind of fiction within prototype software called Mo文oNExplorer. It describes how a...

    Go to contribution page
  39. Andrew Janco (Princeton University)
    9/8/26, 11:20 AM

    This demonstration builds on our experience with 19th-century archival documents from Chocó, Colombia. As researchers digitised these documents, there was a concurrent need for datafication to assess the collection's scope, content, and research potential. We developed a minimal tool to extract text using vision-language models (VLMs), identify and normalise entities, and create a search...

    Go to contribution page
  40. Sebastian Sobecki (University of Toronto), Sam Grieggs (Indiana University of Pennsylvania)
    9/8/26, 11:30 AM

    Communities of Practice is funded by the Social Sciences and Humanities Research Council of Canada and explores how the multilingual urban bureaucracy of London — the world of Chaucer and his contemporaries — shaped late medieval literature. Bringing together linguistic, palaeographical, literary-historical, and computational methods, the project examines how networks of clerks employed in...

    Go to contribution page
  41. Daniel Tubb (Anthropology, University of New Brunswick, Canada)
    9/8/26, 11:40 AM

    In this talk, I demo Fichero. Fichero is a macOS app, a SwiftUI front end over a FastAPI Python engine, that puts machine-learning and AI workflows, using local and hosted models, in researchers' hands. Developed from a collaboration with Andy Janco (Princeton) and Ann Farnsworth-Alvear (UPenn), I demonstrate Fichero on the Circuit Court Archive of Istmina, Colombia. The archive is a British...

    Go to contribution page
  42. Martin Roček (Institute for Medieval Research, Austrian Academy of Sciences and Faculty of Arts, Charles University)
    9/8/26, 12:00 PM

    What if we reframe intertextuality detection as a retrieval task? In this talk, I will take quotations from an author and rank them against corpus of possible sources. This allows me to measure recall@10, precision@1, nDCG@10 and compare four different methods that are commonly used: BM25, dense embeddings (generic and Latin-specific), reciprocal-rank fusion, and reranking by cross-encoder or...

    Go to contribution page
  43. Giuseppe De Gregorio (Computer Vision Center - CVC - Barcelona)
    9/8/26, 12:00 PM

    Handwritten Text Recognition assumes what encrypted historical manuscripts withhold: a known script, and labelled examples of it. Two questions therefore come first. Which script family does this document belong to, and which symbols does it actually use? Both are still answered by hand, an expert-bound process that does not scale to collections of hundreds of manuscripts. This talk reports...

    Go to contribution page
  44. Jessie Dummer (University of Pennsylvania Libraries)
    9/8/26, 1:30 PM

    With the rapid improvement of handwritten text recognition models, large language models, and multimodal AI models, the University of Pennsylvania Libraries is assessing how these technologies can be applied at scale across its handwritten collections. This presentation will focus on an emerging project to apply HTR to a group of fifteenth-century Italian manuscripts written in humanistic...

    Go to contribution page
  45. Michael Schonhardt (TU Darmstadt / Akademie der Wissenschaften und der Literatur | Mainz)
    9/8/26, 1:30 PM

    Automatic Text Recognition has become the most widely adopted application of machine learning in the humanities, yet projects often encounter challenges when integrating existing models into their workflow. This paper dicusses the source of these challenges as a mismatch between requirement profiles - a concept broader than transcription guidelines, constituted by three dimensions:...

    Go to contribution page
  46. Olaf Berg (Ruhr-Universität Bochum)
    9/8/26, 1:50 PM

    My presentation addresses the technical challenges of applying automated text recognition (ATR) to demographic data preserved in late Ottoman census registers. It builds on the experience we gained in the LOOP research project—Late Ottoman Palestinians. Two issues are central. First, the registers are handwritten in Ottoman Turkish, an under-resourced historical language for which effective...

    Go to contribution page
  47. Michael Lužný (National Library of the Czech Republic)
    9/8/26, 1:50 PM

    After several years of development, the transition to the new version of the Manuscriptorium digital library was completed last autumn. The development, however, is still ongoing – the most significant and recent part of which is the TEI-compatible full-text module. As a result, Manuscriptorium offers its users an alternative core that shifts the central perspective from visual representation...

    Go to contribution page
  48. Osama Eshera (University of Maryland)
    9/8/26, 2:10 PM

    We study word-level text-to-image mapping in Arabic-script manuscripts: given a line image and its transcription, locate each word. At corpus scale, accurate word geometry opens new possibilities for digital paleographical analysis while improving ATR both upstream and downstream. This mapping is learned from 25,000 line crops drawn from 87 manuscripts, supervising word geometry only through...

    Go to contribution page
  49. Tim Geelhaar (Goethe Universität Frankfurt am Main)
    9/8/26, 2:10 PM

    The Latin Text Archive (LTA) is an online platform hosted by the Berlin-Brandenburg Academy of Sciences (BBAW) since 2020 (https://LTA.bbaw.de). Its primary objective is to facilitate computer-assisted semantic analysis of Latin texts and corpora spanning various epochs and genres. Conceived as an open platform from its inception in the mid-2000s, it enables external editors and text providers...

    Go to contribution page
  50. Johannes Knüchel (Austrian National Library)
    9/8/26, 2:30 PM

    One of the Austrian National Library’s Strategic Goals for 2023–2027 is to create new ways for users to explore its collections. The improvement of ATR, especially OCR, falls within this context. During my presentation, I will talk about the currently ongoing OCR Project (2024–2027) and one of its main outcomes, an internal OCR Service. This Service contains a modular OCR pipeline using...

    Go to contribution page
  51. Ursula Stampfer (Bayerische Staatsbibliothek)
    9/8/26, 2:30 PM

    The German Manuscript Centres have a long-standing tradition of developing standards, infrastructures, and services for the scholarly description and accessibility of medieval manuscripts. With the transition from printed catalogues to digital infrastructures such as Manuscripta Mediaevalia and, since 2023, the Handschriftenportal, digitisation has fundamentally expanded both the possibilities...

    Go to contribution page
  52. 9/8/26, 3:30 PM
    • WG 1 [Room 1 = HS1]: HTR Technology Development (coordinators: Tobias Hodel)
    • WG 2 [Room 2 = SR5]: Document or Handwriting Classification (coordinators: Maria Konstantinidou, John Pavlopoulos, Paraskevi Platanou)
    • WG 3 [Room 3 = SR4]: Methodological Issues of HTR (coordinators: Anna Michalcová and Jan Odstrčilík)
    • WG 4 [Room 4 = SR3]: Language Challenges (coordinators:...
    Go to contribution page
  53. 9/8/26, 5:15 PM
  54. 9/8/26, 6:45 PM
  55. 9/9/26, 9:00 AM