2ⁿᵈ Exchange Meeting of SCOOP: International Network for Automated Text Recognition of Historical Sources
University of Vienna
2nd Exchange Meeting, Vienna, September 7-9. 2026
SCOOP (Source Codes of the Past) is an international network dedicated to the automatic transcription and analysis of handwritten historical sources. It brings together people from very different corners of the scholarly world: historians, philologists, and palaeographers, archivists and librarians; computer scientists and machine learning researchers; software engineers asf. What unites them is a shared challenge: how to use automatic/handwritten text recognition (ATR/HTR) to unlock the vast written heritage of the past, and how to do it well. (Read more about SCOOP.)
After a first smaller meeting at Princeton in June of 2025 (for more information see https://fsp-text-edition-blog.univie.ac.at/?p=128), the second SCOOP Exchange Meeting is taking place in Vienna on 7–9 September 2026 hosted by the Institute for Medieval Research at the Austrian Academy of Sciences and the Faculty for Cultural Historical Studies at the University of Vienna.
Over three days, it will bring together around a hundred members of the network for keynotes, parallel working group sessions, round tables, demonstrations, and open discussion formats, along with: as well as an internal working day devoted to the future of SCOOP:
September 7–8: Part of the conference open to registered external visitors
September 9: A dedicated internal working session for SCOOP members
September 10 (Associated Event): The OCR/HTR Workshop for Under-represented and Under-resourced Languages organized by Alíz Horváth - If you wish to join or present, please contact Alíz Horváth (HorvathA@ceu.edu).
-
-
9:00 AM
→
10:00 AM
Registration 1h
-
10:00 AM
→
10:15 AM
Plenary: Welcome Room 1
Room 1
-
10:15 AM
→
11:30 AM
Keynote Room 1
Room 1
-
10:15 AM
The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text Recognition 1h 15mSpeaker: Thibault Clérice
-
10:15 AM
-
11:30 AM
→
12:00 PM
Coffee break 30m
-
12:00 PM
→
1:30 PM
Working Group 1: WG1-1 - Technologies & Architectures Room 1
Room 1
-
12:00 PM
Large Scale Approaches: Party 20mSpeaker: Benjamin Kiessling (ALMAnaCH, Inria)
-
12:20 PM
VLM Input – Fine-tuning VLMS 20mSpeaker: William Mattingly
-
12:40 PM
AnandaSky 20mSpeaker: Colin Brisson (EPHE Paris)
-
1:00 PM
Metatools & Technological Agnosticism 20mSpeaker: Andy Stauder (READ COOP)
-
12:00 PM
-
1:30 PM
→
2:30 PM
Lunch 1h
-
2:30 PM
→
4:00 PM
Working Group 4: Language Challenges I Room 3
Room 3
-
2:30 PM
Four Scripts, Five Languages, One Digitalization Challenge 20mSpeaker: Ana Mihaljević (Institute for the Croatian language)
-
2:50 PM
Matchbox: creating a combined recognition model for early medieval Celtic languages and Latin 20mSpeaker: Bernhard Bauer
-
3:10 PM
Promising results for under-resourced languages : use cases on Arabic, Greek and Chinese 20mSpeaker: Baptiste Queuche
-
3:30 PM
Using ATR of Yiddish to study language change 20mSpeaker: Seth Kulick
-
2:30 PM
-
2:30 PM
→
4:00 PM
Working Group 1: Training as a continuous process Room 1
Room 1
-
2:30 PM
Piloting HTR in the Library: Experiments and Groundwork 20mSpeaker: Doug Emery (UPenn)
-
2:50 PM
Multi-Model OCR Arbitration: High-Precision OCR for 19th Century Administrative Mass Sources 20mSpeaker: Wolfgang Goederle (Graz/Passau)
-
3:10 PM
From PDF to mARkdown: A Scalable Arabic OCR Pipeline for expanding the OpenITI Corpus 20mSpeaker: Alicia González Martínez (Universität Hamburg)
-
3:30 PM
Anagnostes - Towards a Transformer-based OCR System for Ancient Greek Papyri 20mSpeakers: Anton Repushko, Elena Chepel (University of Vienna)
-
2:30 PM
-
2:30 PM
→
4:00 PM
Working Group 3: Transcription approaches I (Palaeography in focus) Room 2
Room 2
-
2:30 PM
Computational Paleography through Automatic Text Recognition 30mSpeaker: Benjamin Kiessling (ALMAnaCH, Inria)
-
3:00 PM
Towards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic Manuscripts 30mSpeaker: Sajjad Nikfahm-Khubravan (Roshan Institute for Persian Studies)
-
3:30 PM
Leveraging HTR corpora for morphological and metrological paleographical analysis 30mSpeaker: Malamatenia Vlachou (IRHT/LIGM IPParis)
-
2:30 PM
-
4:00 PM
→
4:30 PM
Coffee break 30m
-
4:30 PM
→
6:00 PM
Working Group 4: Language Challenges II Room 3
Room 3
-
4:30 PM
Dark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval Icelandic 20mSpeaker: Katrín Lísa van der Linde Mikaelsdóttir
-
4:50 PM
Impacts of Script Features on Text Recognition: Experiments with Greek and Arabic 20mSpeaker: Christine Roughan
-
5:10 PM
Syriac Manuscripts That Can Talk: Giving Voice to Under-Resourced Scripts in HTR 30mSpeaker: Ephrem Ishac
-
4:30 PM
-
4:30 PM
→
6:00 PM
Working Group 1: Experimental approaches Room 1
Room 1
-
4:30 PM
Utilization of the Open WebUI platform in Arkindex 20mSpeaker: Michal Racyn (Masaryk University)
-
4:50 PM
Comparing HTR engines with Polyscriptor 20mSpeaker: Achim Rabus
-
5:10 PM
Fichero and the Circuit Court Archive of Istmina: An AI Workflow App for Researchers 20mSpeaker: Daniel Tubb
-
5:30 PM
Document Layout Analysis for Glossed Medieval Manuscripts: Comparing Kraken 6 and YOLO26 20mSpeaker: Tristan Repolusk (University of Graz)
-
4:30 PM
-
4:30 PM
→
6:00 PM
Working Group 3: Transcription approaches II (Transcription decisions, their effect, and related terminology) Room 2
Room 2
-
4:30 PM
Otto Dei gram, rogat vram Clam. From facsimile and descriptive transcriptions to “quasi-diplomatic” and interpretative approaches in ATR/HTR 20mSpeaker: Jan Odstrčilík (Institut für Mittelalterforschung, ÖAW)
-
4:50 PM
Lost in Transcription? Reconciling HTR Output with Old Czech Editorial Standards 20mSpeaker: Anna Michalcová (Czech Academy of Sciences / IMAFO, ÖAW)
-
5:10 PM
HTR of Normalized Latin Texts: Insights from Liturgical Manuscripts 20mSpeaker: Paweł Figurski (Polish Academy of Sciences)
-
5:30 PM
Reading Glagolitic in the Twenty-First Century: Between Philology, HTR, and AI 20mSpeaker: Ana Mihaljević (Institute for the Croatian language)
-
4:30 PM
-
6:00 PM
→
6:15 PM
Break 15m
-
6:15 PM
→
7:30 PM
Keynote Room 1
Room 1
-
6:15 PM
New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and Interpretation 1h 15mSpeakers: Anna Dolganov, David Smith (Northwestern)
-
6:15 PM
-
7:30 PM
→
9:00 PM
Reception 1h 30m
-
9:00 AM
→
10:00 AM
-
-
10:00 AM
→
11:30 AM
Working Group 5: Datasets and Institutions Room 2
Room 2
-
10:00 AM
Pipelines and Workflows 20mSpeaker: Jessie Dummer
-
10:20 AM
Manuscriptorium Full-Text Module – The newest component of the Manuscriptorium digital library 20mSpeaker: Michael Lužný
-
10:40 AM
How do you revive a legacy dataset? 20mSpeaker: Tim Geelhaar
-
11:00 AM
The current status of the field of HTR within the German manuscript centres 20mSpeaker: Ursula Stampfer
-
10:00 AM
-
10:00 AM
→
11:30 AM
Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (talks) Room 3
Room 3
-
10:00 AM
Testing the Reuse Value of Pracalit Ground Truth for Bhujimol Manuscripts 30mSpeaker: Alexander O'Neill
-
10:30 AM
From Old Icelandic HTR Outputs to Normalised Texts through Seq2Seq Transformers 30mSpeaker: Nikola Krisztian Czindrity
-
11:00 AM
New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer 30mSpeaker: Seth Kulick
-
10:00 AM
-
10:00 AM
→
11:30 AM
Working Group 1: Roundtable Room 1
Room 1
-
11:30 AM
→
12:00 PM
Coffee Break 30m
-
12:00 PM
→
1:30 PM
Working Group 4: Roundtable 🎤 - Language Challenges Room 2
Room 2
-
12:00 PM
→
1:30 PM
Working Group 6: Leveraging Outputs: Text Reuse, NLP, and More (demos) Room 3
Room 3
-
12:00 PM
New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer 30mSpeaker: Wayne de Fremery
-
12:30 PM
From Document Images to Research Catalogue 30mSpeaker: Andrew Janco
-
1:00 PM
TBD 30mSpeaker: Martin Roček
-
12:00 PM
-
12:00 PM
→
1:30 PM
Working Group 2: Handwriting Classification I Room 1
Room 1
-
12:00 PM
Classifying Squeezes Again: Initial Results from an ICDAR Contest 30mSpeaker: Aaron Hershkowitz
-
12:30 PM
Communities of Practice: Capturing the Aspect of Late Medieval Handwriting 30mSpeakers: Sam Grieggs, Sebastian Sobecki
-
1:00 PM
Script classification and alphabet identification 30mSpeaker: Giuseppe de Gregorio
-
12:00 PM
-
1:30 PM
→
2:30 PM
Lunch 1h
-
2:30 PM
→
4:00 PM
Working Group 6: Unconference Room 3
Room 3
-
2:30 PM
→
4:00 PM
Working Group 2: Handwriting Classification II Room 1
Room 1
-
2:30 PM
[TBD - Handwriting classification] 30mSpeakers: Asimina Paparrigopoulou, Paraskevi (Vivian) Platanou
-
3:00 PM
Classification of Armenian manuscripts without OCR/HTR preprocessing 30mSpeaker: Tara Andrews
-
3:30 PM
[TBD - Explainability] 30mSpeaker: Serena Ammirati
-
2:30 PM
-
2:30 PM
→
4:00 PM
Working Group 3: Building ATR/HTR pipelines Room 2
Room 2
-
2:30 PM
The Pragmatics of ATR Pipelines: Structural Dilemmas in Designing Workflows for Historical Corpora 20mSpeaker: Michael Schonhardt (Univ. Freiburg)
-
2:50 PM
Building an ATR pipeline for tabular data and script translation 20mSpeaker: Olaf Berg (Ruhr-Universität Bochum, AI:RUB, ERC-LOOP)
-
3:10 PM
Iterative HTR: A new pipeline for automatic text recognition of Arabic-script 20mSpeaker: Osama Eshera (University of Maryland)
-
3:30 PM
Developing a pipeline for OCR at the Austrian National Library 20mSpeaker: Johannes Knüchel (Austrian National Library)
-
2:30 PM
-
4:00 PM
→
4:30 PM
Coffee Break 30m
-
4:30 PM
→
6:00 PM
Working Groups working on conclusions 1h 30m
-
6:00 PM
→
6:15 PM
Short Break 15m
-
6:15 PM
→
7:45 PM
Final Roundtable - Open data, open code, open minds in AI 1h 30m
-
7:45 PM
→
8:00 PM
Conclusion of the public part 15m
-
10:00 AM
→
11:30 AM
-
-
10:00 AM
→
6:05 PM
Internal meeting of SCOOP 8h 5m
-
10:00 AM
→
6:05 PM