From Document Images to Research Catalogue

Sep 8, 2026, 11:20 AM
20m
Room 3

Room 3

Speaker

Andrew Janco (Princeton University)

Description

This demonstration builds on our experience with 19th-century archival documents from Chocó, Colombia. As researchers digitised these documents, there was a concurrent need for datafication to assess the collection's scope, content, and research potential. We developed a minimal tool to extract text using vision-language models (VLMs), identify and normalise entities, and create a search interface. This catalogue tool, named ficherito, addresses a common use case in which researchers need machine-readable text and structured data for exploratory data analysis. This pilot identified the need for more full-featured software, called fichero, demonstrated by Daniel Tubb.

Presentation materials

There are no materials yet.