SCOOP: Source Codes of the Past

Source Codes of the PastInternational Network for Automated Text Recognition of Historical Sources

SCOOP brings together historians, philologists and palaeographers, archivists and librarians, computer scientists and software engineers around a shared challenge: making the handwritten heritage of the past machine-readable, across scripts, languages and centuries.

The network runs six working groups, an open exchange platform, and regular in-person meetings — most recently in Vienna in September 2026, hosted by the Austrian Academy of Sciences and the University of Vienna.

SCOOP logo

About

SCOOP (Source Codes of the Past) is an international network dedicated to the automatic transcription and analysis of handwritten historical sources. It brings together people from very different corners of the scholarly world: historians, philologists, and palaeographers; archivists and librarians; computer scientists, machine learning researchers, and software engineers. What unites them is a shared challenge: how to use automatic/handwritten text recognition (ATR/HTR) to unlock the vast written heritage of the past, and how to do it well.

Recent years have brought remarkable progress in text recognition technologies, but much work remains in adapting them to the diversity of historical materials, ranging from different scripts and writing systems, varied textual traditions and manuscript structures, to low-resource languages for which training data and tools are scarce. At the same time, ATR is increasingly interwoven with other methodologies, from dataset curation and text reuse analysis to the automation of editorial workflows. These developments are transforming how researchers engage with manuscript sources, and they raise questions that no single discipline can answer alone.

SCOOP exists to bring these conversations together. Technological development, methodological reflection, and the practical needs of editors, cataloguers, and collecting institutions too often happen in separate communities; the network's aim is to let technological and humanistic expertise inform one another directly, across the boundaries of disciplines, languages, and scripts.

How the network works

The work of SCOOP is organised into six working groups, each covering different focus areas. These working groups structure both the meetings and the exchange platform, which supports preparation for the in-person exchange meetings as well as the ongoing discussions and cooperation between the various groups and participants. Each member of SCOOP is a member of one or more working groups.

WG1
HTR Technology Development
WG2
Document or Handwriting Classification
WG3
Methodological Issues of HTR
WG4
Language Challenges
WG5
Datasets and Institutions
WG6
Leveraging Outputs: Text Reuse, NLP and More

Join the network — and our Discord

SCOOP is open to anyone working on the automatic transcription and analysis of historical sources, across disciplines, scripts, and languages. If you'd like to join the network, write to scoop@oeaw.ac.at, and put Anna Michalcová (a.michalcova@ujc.cas.cz) and Jan Odstrčilík (jan.odstrcilik@oeaw.ac.at) into CC. You will be included into the future communication and given access to the Discord server.

Discord is where much of SCOOP's exchange happens between meetings. It's a free chat platform where conversations are organised into channels: each working group has its own space to discuss its topics, share materials or simply to get to know each other. It's also the easiest place to ask a quick question, share a new tool, dataset, or paper, announce a call or event, or find a collaborator who has already struggled with the same script, language, or pipeline you're facing now.

If you're new to the network, the server is the best way to get a feel for what's going on in SCOOP and to introduce yourself and your project before meeting everyone in person.

Organisers of SCOOP network

  • Tara Andrews

    University of Vienna

    • Scientific committee
  • Gerda Heydemann

    Friedrich-Meinecke-Institut für Geschichte und Historische Kulturwissenschaften

    • Scientific committee
  • Tobias Hodel

    Digital Humanities, University of Bern

    • Scientific committee
  • Alíz Horváth

    Central European University

    • Scientific committee
  • Maria Konstantinidou

    Democritus University of Thrace

    • Scientific committee
  • Anna Michalcová

    Czech Language Institute, Czech Academy of Sciences / Institute for Medieval Research, Austrian Academy of Sciences

    • Scientific committee
  • Jan Odstrčilík

    Institute for Medieval Research, Austrian Academy of Sciences

    • Scientific committee
  • John Pavlopoulos

    Athens University of Economics and Business, and Archimedes, Athena Research Center

    • Scientific committee
  • Paraskevi Platanou

    Athens University of Economics and Business, and Archimedes, Athena Research Center

    • Scientific committee
  • Helmut Reimitz

    Institute for Medieval Research, Austrian Academy of Sciences / Institute for Austrian History Research, University of Vienna

    • Scientific committee
  • Martin Roček

    Institute for Medieval Research, Austrian Academy of Sciences / Faculty of Arts, Charles University

    • Scientific committee
  • Christine Roughan

    Center for Digital Humanities / MARBAS, Princeton University

    • Scientific committee
  • Sofia Torallas Tovar

    Institute for Advanced Study, Princeton

    • Scientific committee
  • Lucia Waldschuetz

    Princeton University / Institute for Medieval Research, Austrian Academy of Sciences

    • Scientific committee
  • Thomas Wallnig

    University of Vienna

    • Scientific committee

Organising institutions