Search Results
Discussion Paper
Combining AI and Established Methods for Historical Document Analysis
This paper examines methodological approaches for extracting structured data from large-scale historical document archives, comparing “hyperspecialized” versus “adaptive modular” strategies. Using 56 years of Philadelphia property deeds as a case study, we show the benefits of the adaptive modular approach leveraging optical character recognition (OCR), full-text search, and frontier large language models (LLMs) to identify deeds containing specific restrictive use language— achieving 98% precision and 90–98% recall. Our adaptive modular methodology enables analysis of ...