SCOOP: Source Codes of the Past

Back to the programme

From Document Images to Research Catalogue

Andrew Janco ・ Princeton University

Tuesday, 8th September ・ 11:00 - 12:30 ・ WG6-2 ・ Leveraging Outputs: Text Reuse, NLP, and More (demos) ・ Room 3

This demonstration builds on our experience with 19th-century archival documents from Chocó, Colombia. As researchers digitised these documents, there was a concurrent need for datafication to assess the collection's scope, content, and research potential. We developed a minimal tool to extract text using vision-language models (VLMs), identify and normalise entities, and create a search interface. This catalogue tool, named ficherito, addresses a common use case in which researchers need machine-readable text and structured data for exploratory data analysis. This pilot identified the need for more full-featured software, called fichero, demonstrated by Daniel Tubb.