|
Having trouble viewing this email?
Click here.
|
|
SCOOP: Source Codes of the Past
|
|
|
The Programme Is Out
|
|
2nd Exchange Meeting of SCOOP — and an invitation to join the keynotes online
|
September 7–9, 2026, University of Vienna, Währingerstraße 29, 1090 Vienna
|
Hosted by the Austrian Academy of Sciences (Institute for Medieval Research) and the University of Vienna (Faculty of Historical and Cultural Studies, Institute for Austrian Historical Research)
|
|
|
Join us online
|
|
If you cannot make it to Vienna, you are still very welcome to take part. Both keynotes on Monday, 7 September, are streamed on Zoom and open to everyone — no registration needed, just follow the link at the time of the talk. The parallel working group sessions are not streamed; they take place on site only.
|
|
Keynote · Monday, 7 September, 09:15–10:30 CEST
|
|
Thibault Clérice · ALMAnaCH, Inria
|
|
The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text Recognition
|
|
Join this keynote on Zoom
|
|
|
Keynote · Monday, 7 September, 17:15–18:30 CEST
|
|
Anna Dolganov & David Smith · Austrian Academy of Sciences, Northeastern University
|
|
New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and Interpretation
|
|
Join this keynote on Zoom
|
|
|
All times are Central European Summer Time (CEST, UTC+2). The Zoom links are also published next to each keynote on the programme page, so you can always find the current one there.
|
|
|
Programme
|
|
All three days, with the speakers and papers of every working group session. The programme is also kept up to date on the website, which stays in sync with Indico.
|
|
Monday 7th September
|
|
08:00 – 09:00 · Registration
|
|
|
09:00 – 09:15 · Plenary
|
|
Welcome
|
|
|
09:15 – 10:30 · Keynote · streamed
|
|
Thibault Clérice · ALMAnaCH, Inria
|
|
The Transcription/Edition Distinction as a Foundation for Trustworthy Automatic Text Recognition
|
|
Join on Zoom
|
|
|
10:30 – 11:00 · Coffee break
|
|
|
11:00 – 12:30 · Session
|
|
WG1-1 Technologies & Architectures · Room 1
|
William Mattingly · Yale University
Finetuning VLMs in a World of Zero-Shot Frontier Models: A Paradigm Shift toward Knowledge Distillation and Local Deployment
|
Colin Brisson · Ecolé pratique des hautes études
AnandaSky
|
Andy Stauder · READ COOP
Metatools & Technological Agnosticism
|
|
|
|
|
|
13:30 – 15:00 · Parallel sessions
|
|
WG1-2 Training as a continuous process · Room 1
|
Doug Emery · University of Pennsylvania Libraries
Piloting HTR in the Library: Experiments and Groundwork
|
Wolfgang Göderle · University of Graz | University of Passau | MPI GEA
Multi-Model OCR Arbitration: High-Precision OCR for 19th Century Administrative Mass Sources
|
Alicia González Martínez · Hamburg University
From PDF to mARkdown: A Scalable Arabic OCR Pipeline for expanding the OpenITI Corpus
|
Elena Chepel & Anton Repushko · University of Vienna; Independent Researcher
Anagnostes — Towards a Transformer-based OCR System for Ancient Greek Papyri
|
|
|
WG3-1 Transcription approaches I (Palaeography in focus) · Room 2
|
Benjamin Kiessling · ALMAnaCH, Inria Paris
Computational Paleography through Automatic Text Recognition
|
Sajjad Nikfahm-Khubravan · Roshan Institute for Persian Studies, University of Maryland
Towards Geometry-Based Scribal Hand Analysis: A Word-Level Graphometric Framework for Arabic Manuscripts
|
Malamatenia Vlachou · IRHT/CNRS-ENPC
Leveraging HTR for Script Characterization: Toward a Unified Computational Framework for Palaeographical Analysis
|
Dominique Stutzmann · IRHT-CNRS / HU Berlin
Workflows and granularities: text, image, and what remains
|
|
|
WG4-1 Language Challenges I · Room 3
|
Ana Mihaljević · Institute for the Croatian language
Four Scripts, Five Languages, One Digitalization Challenge
|
Bernhard Bauer · University of Graz
Matchbox: Creating a Combined Recognition Model for Early Medieval Celtic Languages and Latin
|
Baptiste Queuche · Calfa
How hybrid HTR+VLM approaches are unlocking the processing of under-resourced, non-western languages
|
Seth Kulick · Linguistic Data Consortium, University of Pennsylvania
Using OCR to expand historical treebanks
|
Andrew Janco · Princeton University
Which Languages Are In-Vocabulary?
|
Isabelle Marthot-Santaniello · University of Basel
The language of Greek papyri? One millennium of writing and its complexity
|
|
|
|
15:00 – 15:30 · Coffee break
|
|
|
15:30 – 17:00 · Parallel sessions
|
|
WG1-3 Experimental approaches · Room 1
|
Michal Racyn · Masaryk University
Utilization of the Open WebUI platform in Arkindex
|
Tristan Repolusk · Department of Digital Humanities, University of Graz
Document Layout Analysis for Glossed Medieval Manuscripts: Comparing Kraken 6 and YOLO26
|
Elise Wang · California State University, Fullerton
Finding our way through the National Archives: HTR for indexing a large corpus
|
Tobias Hodel · University of Bern
Insights from an Unsound Experiment: Testing Kraken, TrOCR, and VLMs with an LLM Judge
|
|
|
WG3-2 Transcription approaches II (Transcription decisions, their effect, and related terminology) · Room 2
|
Jan Odstrčilík · Institute for Medieval Research, Austrian Academy of Sciences
Otto Dei gram, rogat vram Clam. From facsimile and descriptive transcriptions to “quasi-diplomatic” and interpretative approaches in ATR/HTR
|
Anna Michalcová · Czech Language Institute, Czech Academy of Sciences; Institute for Medieval Research, Austrian Academy of Sciences
Lost in Transcription? Reconciling HTR Output with Old Czech Editorial Standards
|
Ana Mihaljević · Institute for the Croatian language
Reading Glagolitic in the Twenty-First Century: Between Philology, HTR, and AI
|
Paweł Figurski · Polish Academy of Sciences
HTR of Normalized Latin Texts: Insights from Liturgical Manuscripts
|
|
|
WG4-2 Language Challenges II · Room 3
|
Katrín Lísa van der Linde Mikaelsdóttir · University of Iceland
Dark Vellum, Dense Diacritics: Challenges in Training HTR Models for Medieval Icelandic
|
Christine Roughan · Princeton University
Impacts of Script Features on Text Recognition: Experiments with Greek and Arabic
|
Ephrem Aboud Ishac · Institute for Medieval Research, Austrian Academy of Sciences
Syriac Manuscripts That Can Talk: Giving Voice to Under-Resourced Scripts in HTR
|
|
|
|
|
|
17:15 – 18:30 · Keynote · streamed
|
|
Anna Dolganov & David Smith · Austrian Academy of Sciences, Northeastern University
|
|
New Epistemic Frontiers: LLMs for Linking Transcription, Editing, and Interpretation
|
|
Join on Zoom
|
|
|
18:30 – 20:00 · Reception
|
|
|
Tuesday 8th September
|
|
09:00 – 10:30 · Parallel sessions
|
|
WG1-4 Roundtable · Room 1
|
|
WG2-1 Handwriting Classification I · Room 2
|
Asimina Paparrigopoulou & Paraskevi Platanou · Democritus University of Thrace; Athens University of Economics and Business, and Archimedes, Athena Research Center
From Handwriting Classification Errors to Paleographic Evidence: Graphic Compensation in Ancient Greek Documentary Hands
|
Tara Andrews · University of Vienna
Classification of Armenian manuscripts without OCR/HTR preprocessing
|
Serena Ammirati & Paolo Merialdo · Università degli Studi Roma Tre
Beyond the Heatmap: What Faithful Explanations Can Do for Handwriting Identification
|
|
|
WG6-1 Leveraging Outputs: Text Reuse, NLP, and More (talks) · Room 3
|
Alexander O’Neill · Musashino University
Testing the Reuse Value of Pracalit Ground Truth for Bhujimol Manuscripts
|
Nikola Krisztian Czindrity · University of Vienna
From Old Icelandic HTR Outputs to Normalised Texts through Seq2Seq Transformers
|
Seth Kulick · Linguistic Data Consortium, University of Pennsylvania
Using OCR to expand a treebank of historical Yiddish: A Progress Report
|
|
|
|
10:30 – 11:00 · Coffee break
|
|
|
11:00 – 12:30 · Parallel sessions
|
|
WG4-3 Roundtable — Language Challenges · Room 1
|
|
WG2-2 Handwriting Classification II · Room 2
|
Aaron Hershkowitz · The Institute for Advanced Study
Classifying Squeezes Again: Initial Results from an ICDAR Contest
|
Sebastian Sobecki & Sam Grieggs · University of Toronto; Indiana University of Pennsylvania
Communities of Practice: Capturing the Aspect of Late Medieval Handwriting
|
Giuseppe De Gregorio · Computer Vision Center — CVC, Barcelona
Script classification and alphabet identification
|
|
|
WG6-2 Leveraging Outputs: Text Reuse, NLP, and More (demos) · Room 3
|
Wayne de Fremery · Dominican University of California
New Ways to Transform and Explore Older Korean Texts: From ATR to Reading Texts to Agentic Retrieval in Mo文oNExplorer
|
Andrew Janco · Princeton University
From Document Images to Research Catalogue
|
Daniel Tubb · Anthropology, University of New Brunswick
Fichero and the Circuit Court Archive of Istmina: An AI Workflow App for Researchers
|
Martin Roček · Institute for Medieval Research, Austrian Academy of Sciences; Faculty of Arts, Charles University
Intertextuality as a Retrieval Task: Benchmarking Text Reuse in Classical and Medieval Latin
|
|
|
|
|
|
13:30 – 15:00 · Parallel sessions
|
|
WG5-1 Datasets and Institutions · Room 1
|
Jessie Dummer · University of Pennsylvania Libraries
Integrating HTR into Digital Library Workflows
|
Michael Lužný · National Library of the Czech Republic
Manuscriptorium Full-Text Module: Building a Digital-Edition Infrastructure
|
Tim Geelhaar · Goethe Universität Frankfurt am Main
Reviving the Legacy: How to Adapt the Latin Text Archive for the Age of AI
|
Ursula Stampfer · Bayerische Staatsbibliothek
HTR in the German Manuscript Centres: Current Status, Challenges, and Perspectives
|
|
|
WG3-3 Building ATR/HTR pipelines · Room 2
|
Michael Schonhardt · TU Darmstadt / Akademie der Wissenschaften und der Literatur, Mainz
The Pragmatics of ATR Pipelines: Structural Dilemmas in Designing Workflows for Historical Corpora
|
Olaf Berg · Ruhr-Universität Bochum
Building an ATR pipeline for tabular data and script translation
|
Osama Eshera · University of Maryland
Iterative HTR: A new pipeline for automatic text recognition of Arabic-script
|
Johannes Knüchel · Austrian National Library
Building an OCR Pipeline at the Austrian National Library
|
|
|
WG6-3 Unconference · Room 3
|
|
|
15:00 – 15:30 · Coffee break
|
|
|
15:30 – 17:00 · Session
|
|
Working Groups working on conclusions
|
|
|
17:00 – 17:15 · Short break
|
|
|
17:15 – 18:45 · Roundtable
|
|
Final Roundtable — Open data, open code, open minds in AI
|
|
|
18:45 – 19:00 · Conclusion
|
|
Conclusion of the public part
|
|
|
Wednesday 9th September
|
|
09:00 – 17:00 · Meeting
|
|
Internal meeting of SCOOP
|
|
For network members only.
|
|
|
On-site places are still available — register via the Indico registration page. If you have any questions, please contact: jan.odstrcilik@oeaw.ac.at; helmut.reimitz@oeaw.ac.at; Thomas.Wallnig@univie.ac.at
|
|
|
Sponsored by
|
Austrian Academy of Sciences
Center for Digital Humanities, Princeton University
Faculty of Historical and Cultural Studies, University of Vienna
Humanities Initiative, Princeton University
Institute for Austrian Historical Research, University of Vienna
Institute for Medieval Research, Austrian Academy of Sciences
Machine Learning Topical platform (MLA2S), Austrian Academy of Sciences
Princeton Athens Center, Princeton University
School of Historical Studies, Institute for Advanced Study
The Seeger Center for Hellenic Studies, Princeton University
University of Vienna
|
|
|