PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026Journal of Open Humanities Data0 citationsOpen Access

Are We There Yet? Notes Towards Benchmarking an Experimental AI-Assisted Workflow for Humanities Data Cleaning and Reconciliation

EMErin McCarthy

Key Points

  • This paper aims to discuss the development of an AI-assisted data cleaning pipeline for a humanities project.
  • Developed AI-assisted data cleaning pipeline for early modern poetry data.
  • Aggregated and reconciled datasets about early modern verse circulation.
  • Created unique identifiers and authorities to manage data accurately.
  • Identified challenges and bottlenecks in the pipeline process.
  • Recorded observations about the AI-assisted workflow.
  • Highlighted the need for benchmarks in computational methods for digital humanities.

Abstract

This paper introduces a novel AI-assisted pipeline developed to prepare data for the European Research Council-funded project “STEMMA: Systems of Transmitting Early Modern Manuscript Verse, 1475–1700.” Now approaching its midpoint, STEMMA develops and applies a data-driven approach to provide the first comprehensive study of the circulation of early modern English poetry in manuscript. The project began by aggregating and reconciling five of the largest and most authoritative existing datasets about early modern verse circulation. The sheer volume of data, along with the need to preserve early modern English spelling and scribal idiosyncrasies for later analyses, meant that off-the-shelf data cleaning tools like OpenRefine were not fit for purpose. To that end, our software developer created a staged pipeline to aid the removal of duplicates, creation of authorities, reconciliation, and assignment of unique identifiers. The rapid and pragmatic way that this process was developed and deployed means that we did not take the time to benchmark it, nor is it feasible to do so retrospectively. However, this discussion paper records observations from this process and reflects on challenges and bottlenecks as well as opportunities. It points the way toward future benchmarks that are increasingly needed for novel applications of computational methods in the digital humanities. It also briefly considers the relationship between technical benchmarking and research project management.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Erin McCarthy (2026) studied this question.

synapsesocial.com/papers/69be369a6e48c4981c675a67https://doi.org/10.5334/johd.490
Ask AI
Helpful
Bookmark
Share
View Full Paper