The increasing demand for data related to life cycle assessment, such as environmental product declarations, necessitates the development of more efficient approaches for the preparation and integration of life cycle assessment-relevant data. A central challenge in this context is data mapping, specifically the consistent linkage between foreground and background datasets. Of particular importance is the identification of so-called module-sensitive components, such as printed circuit boards, which consist of highly interdependent materials in terms of their environmental impact. The environmental impacts of these components are often underestimated by conventional approaches and typically require manual identification. To address this challenge, a machine learning-based framework is proposed to enable the automated classification of such components. To the best of the authors’ knowledge, this work presents the first framework that systematically combines substance-based information with textual similarity features for automated identification of module-sensitive components in life cycle inventory data preparation. The framework outlines how a sufficiently robust classification model can be prepared, optimized, and applied. It was evaluated using a dataset of more than 13,000 individual components from the automotive industry. The results demonstrate that combining substance-related information with name similarity metrics significantly improves classification accuracy. Among the tested algorithms, Random Forest achieved the best performance, reaching a classification accuracy of 97% for electronic components such as printed circuit boards, whereas Support Vector Machines and k-Nearest Neighbors yielded substantially lower accuracies. Overall, the proposed framework reduces manual effort, enhances data consistency, and supports the operationalization and scalability of life cycle inventory and, consequently, life cycle assessment preparation processes.
Sturm et al. (2026) studied this question.