2ⁿᵈ Exchange Meeting of SCOOP: International Network for Automated Text Recognition of Historical Sources

Europe/Vienna
University of Vienna

University of Vienna

Währinger Straße 29 1090 Vienna Austria
Description

2nd Exchange Meeting, Vienna, September 7-9. 2026 

SCOOP (Source Codes of the Past) is an international network dedicated to the automatic transcription and analysis of handwritten historical sources. It brings together people from very different corners of the scholarly world: historians, philologists, and palaeographers, archivists and librarians; computer scientists and machine learning researchers; software engineers asf. What unites them is a shared challenge: how to use automatic/handwritten text recognition (ATR/HTR) to unlock the vast written heritage of the past, and how to do it well. (Read more about SCOOP.)

After a first smaller meeting at Princeton in June of 2025 (for more information see https://fsp-text-edition-blog.univie.ac.at/?p=128), the second SCOOP Exchange Meeting is taking place in Vienna on 7–9 September 2026 hosted by the Institute for Medieval Research at the Austrian Academy of Sciences and the Faculty for Cultural Historical Studies at the University of Vienna.

Over three days, it will bring together around a hundred members of the network for keynotes, parallel working group sessions, round tables, demonstrations, and open discussion formats, along with: as well as an internal working day devoted to the future of SCOOP:

September 7–8: Part of the conference open to registered external visitors

September 9: A dedicated internal working session for SCOOP members

September 10 (Associated Event): The OCR/HTR Workshop for Under-represented and Under-resourced Languages organized by Alíz Horváth - If you wish to join or present, please contact Alíz Horváth (HorvathA@ceu.edu).

Registration
Register to 2nd Exchange Meeting of SCOOP (7-8 September)
    • 9:00 AM
      Registration
    • Plenary: Welcome Room 1

      Room 1

    • Keynote Room 1

      Room 1

      • 1
        The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text Recognition
        Speaker: Thibault Clérice
    • 11:30 AM
      Coffee break
    • Working Group 1: WG1-1 - Technologies & Architectures Room 1

      Room 1

      • 2
        Large Scale Approaches: Party
        Speaker: Benjamin Kiessling (ALMAnaCH, Inria)
      • 3
        VLM Input – Fine-tuning VLMS
        Speaker: William Mattingly
      • 4
        AnandaSky
        Speaker: Colin Brisson (EPHE Paris)
      • 5
        Metatools & Technological Agnosticism
        Speaker: Andy Stauder (READ COOP)
    • 1:30 PM
      Lunch
    • Working Group 4: Language Challenges I Room 3

      Room 3

      • 6
        Four Scripts, Five Languages, One Digitalization Challenge
        Speaker: Ana Mihaljević (Institute for the Croatian language)
      • 7
        Matchbox: creating a combined recognition model for early medieval Celtic languages and Latin
        Speaker: Bernhard Bauer
      • 8
        Promising results for under-resourced languages : use cases on Arabic, Greek and Chinese
        Speaker: Baptiste Queuche
      • 9
        Using ATR of Yiddish to study language change
        Speaker: Seth Kulick
    • Working Group 1: Training as a continuous process Room 1

      Room 1

      • 10
        Piloting HTR in the Library: Experiments and Groundwork
        Speaker: Doug Emery (UPenn)
      • 11
        Multi-Model OCR Arbitration: High-Precision OCR for 19th Century Administrative Mass Sources
        Speaker: Wolfgang Goederle (Graz/Passau)
      • 12
        From PDF to mARkdown: A Scalable Arabic OCR Pipeline for expanding the OpenITI Corpus
        Speaker: Alicia González Martínez (Universität Hamburg)
      • 13
        Anagnostes - Towards a Transformer-based OCR System for Ancient Greek Papyri
        Speakers: Anton Repushko, Elena Chepel (University of Vienna)
    • Working Group 3: Transcription approaches I (Palaeography in focus) Room 2

      Room 2

      • 14
        Computational Paleography through Automatic Text Recognition
        Speaker: Benjamin Kiessling (ALMAnaCH, Inria)
      • 15
        Towards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic Manuscripts
        Speaker: Sajjad Nikfahm-Khubravan (Roshan Institute for Persian Studies)
      • 16
        Leveraging HTR corpora for morphological and metrological paleographical analysis
        Speaker: Malamatenia Vlachou (IRHT/LIGM IPParis)
    • 4:00 PM
      Coffee break
    • Working Group 4: Language Challenges II Room 3

      Room 3

      • 17
        Dark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval Icelandic
        Speaker: Katrín Lísa van der Linde Mikaelsdóttir
      • 18
        Impacts of Script Features on Text Recognition: Experiments with Greek and Arabic
        Speaker: Christine Roughan
      • 19
        Syriac Manuscripts That Can Talk: Giving Voice to Under-Resourced Scripts in HTR
        Speaker: Ephrem Ishac
    • Working Group 1: Experimental approaches Room 1

      Room 1

      • 20
        Utilization of the Open WebUI platform in Arkindex
        Speaker: Michal Racyn (Masaryk University)
      • 21
        Comparing HTR engines with Polyscriptor
        Speaker: Achim Rabus
      • 22
        Fichero and the Circuit Court Archive of Istmina: An AI Workflow App for Researchers
        Speaker: Daniel Tubb
      • 23
        Document Layout Analysis for Glossed Medieval Manuscripts: Comparing Kraken 6 and YOLO26
        Speaker: Tristan Repolusk (University of Graz)
    • Working Group 3: Transcription approaches II (Transcription decisions, their effect, and related terminology) Room 2

      Room 2

      • 24
        Otto Dei gram, rogat vram Clam. From facsimile and descriptive transcriptions to “quasi-diplomatic” and interpretative approaches in ATR/HTR
        Speaker: Jan Odstrčilík (Institut für Mittelalterforschung, ÖAW)
      • 25
        Lost in Transcription? Reconciling HTR Output with Old Czech Editorial Standards
        Speaker: Anna Michalcová (Czech Academy of Sciences / IMAFO, ÖAW)
      • 26
        HTR of Normalized Latin Texts: Insights from Liturgical Manuscripts
        Speaker: Paweł Figurski (Polish Academy of Sciences)
      • 27
        Reading Glagolitic in the Twenty-First Century: Between Philology, HTR, and AI
        Speaker: Ana Mihaljević (Institute for the Croatian language)
    • 6:00 PM
      Break
    • Keynote Room 1

      Room 1

      • 28
        New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and Interpretation
        Speakers: Anna Dolganov, David Smith (Northwestern)
    • 7:30 PM
      Reception
    • Working Group 5: Datasets and Institutions Room 2

      Room 2

      • 29
        Pipelines and Workflows
        Speaker: Jessie Dummer
      • 30
        Manuscriptorium Full-Text Module – The newest component of the Manuscriptorium digital library
        Speaker: Michael Lužný
      • 31
        How do you revive a legacy dataset?
        Speaker: Tim Geelhaar
      • 32
        The current status of the field of HTR within the German manuscript centres
        Speaker: Ursula Stampfer
    • Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (talks) Room 3

      Room 3

      • 33
        Testing the Reuse Value of Pracalit Ground Truth for Bhujimol Manuscripts
        Speaker: Alexander O'Neill
      • 34
        From Old Icelandic HTR Outputs to Normalised Texts through Seq2Seq Transformers
        Speaker: Nikola Krisztian Czindrity
      • 35
        New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer
        Speaker: Seth Kulick
    • Working Group 1: Roundtable Room 1

      Room 1

    • 11:30 AM
      Coffee Break
    • Working Group 4: Roundtable 🎤 - Language Challenges Room 2

      Room 2

    • Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (demos) Room 3

      Room 3

      • 36
        New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer
        Speaker: Wayne de Fremery
      • 37
        From Document Images to Research Catalogue
        Speaker: Andrew Janco
      • 38
        TBD
        Speaker: Martin Roček
    • Working Group 2: Handwriting Classification I Room 1

      Room 1

      • 39
        Classifying Squeezes Again: Initial Results from an ICDAR Contest
        Speaker: Aaron Hershkowitz
      • 40
        Communities of Practice: Capturing the Aspect of Late Medieval Handwriting
        Speakers: Sam Grieggs, Sebastian Sobecki
      • 41
        Script classification and alphabet identification
        Speaker: Giuseppe de Gregorio
    • 1:30 PM
      Lunch
    • Working Group 6: Unconference Room 3

      Room 3

    • Working Group 2: Handwriting Classification II Room 1

      Room 1

      • 42
        [TBD - Handwriting classification]
        Speakers: Asimina Paparrigopoulou, Paraskevi (Vivian) Platanou
      • 43
        Classification of Armenian manuscripts without OCR/HTR preprocessing
        Speaker: Tara Andrews
      • 44
        [TBD - Explainability]
        Speaker: Serena Ammirati
    • Working Group 3: Building ATR/HTR pipelines Room 2

      Room 2

      • 45
        The Pragmatics of ATR Pipelines: Structural Dilemmas in Designing Workflows for Historical Corpora
        Speaker: Michael Schonhardt (Univ. Freiburg)
      • 46
        Building an ATR pipeline for tabular data and script translation
        Speaker: Olaf Berg (Ruhr-Universität Bochum, AI:RUB, ERC-LOOP)
      • 47
        Iterative HTR: A new pipeline for automatic text recognition of Arabic-script
        Speaker: Osama Eshera (University of Maryland)
      • 48
        Developing a pipeline for OCR at the Austrian National Library
        Speaker: Johannes Knüchel (Austrian National Library)
    • 4:00 PM
      Coffee Break
    • 49
      Working Groups working on conclusions
    • 6:00 PM
      Short Break
    • 50
      Final Roundtable - Open data, open code, open minds in AI
    • 51
      Conclusion of the public part
    • 52
      Internal meeting of SCOOP