From Documents to Data
Structuring the Historical Record at Scale
Historical documents contain vast amounts of information that has traditionally been difficult to process en masse. Archives want to be able to better describe their holdings and make them more discoverable; researchers wish to conduct large scale analysis of historical data; while genealogists are interested in connecting different individuals together, and understanding the course of their lives.
This roundtable will examine the challenges that these different types of projects have in creating structured information from unstructured documents at scale. The role of human expertise, both in supervising these pipelines and in thinking about the structures that most usefully describe the documents. Participants will discuss their experiences in automating data extraction, whether that is manual work, algorithmic approaches, traditional deep learning, or Large Language Models. Together, the panelists will share insights from different perspectives on the practical workings of these projects, exploring the many potential uses of these technologies.