As part of the broader initiative “Künstliche Intelligenz-gestützte Infrastruktur zur Zeitzeugen-, Erinnerungs- und Quellenforschung” (KIZEQ) at the University of Trier, KIZEQ/CORE develops a curated, semantically annotated and FAIR/CARE-compliant data corpus based on approximately 10,000 digitised reparation case files from the Amt für Wiedergutmachung (AfW) in Saarburg. Between 1953 and the late 1980s, the AfW processed more than 940,000 compensation claims submitted by victims of National Socialist persecution, primarily from outside Europe. These highly sensitive files contain autobiographical statements, expert reports, legal assessments and extensive administrative correspondence. Taken together, they constitute one of the most semantically complex and ethically sensitive archival corpora of postwar Europe. The project addresses both pillars of the DFG funding line “Data Corpora for Artificial Intelligence”. First, it creates a semantically enriched training corpus for AI systems operating in historically embedded and ethically sensitive research contexts. Second, it contributes to the development of sustainable data infrastructures aligned with NFDI4Memory, CLARIN-D and Text+, ensuring long-term interoperability, reuse and scholarly integration. KIZEQ/CORE provides a structured, multi-format corpus designed to support several core use cases in computational history and digital hermeneutics. Corpus-based narrative analysis enables AI systems to identify and reconstruct fragmented life histories shaped by trauma, embedded in bureaucratic and testimonial materials. Ontological modelling extracts legal, medical and institutional entities and integrates them into semantic graphs and knowledge infrastructures. Responsible anonymisation procedures are developed through role-based abstraction, reversible pseudonymisation and normative filtering, enabling privacy-preserving forms of data publication. Finally, explainable AI applications are supported by training models on heterogeneous, multilingual and multi-register historical data in order to enhance transparency and accountability in algorithmic interpretation. The corpus is modular in structure and encoded in TEI/XML and JSON-LD, accompanied by CMDI and PROV-O metadata. Access to the data will be enabled through open standards and interoperable APIs, including publication on platforms such as CLARIN-D, Zenodo and the Trier Research Data Repository (FRET, ViDa). KIZEQ/CORE constitutes the first phase of a long-term digital infrastructure strategy aimed at making the AfW archive machine-readable, semantically explorable and ethically interpretable. In this way, the project establishes a bridge between analogue documentation and AI-assisted analysis in the “post-witness era”. At the same time, it provides a foundational corpus for future integration with related collections such as restitution archives, oral history holdings and postcolonial reparation files, thereby contributing to the development of new standards for responsible digital source research.
Massimiliano Livi (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: