PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 19, 2026Journal of Cleaner Production1 citationsOpen Access

Automated clustering of foreground data for effective dataset integration in life cycle impact assessment

View Full Paper
JSJohannes H.L. SturmTNTill Justus NiemannJHJohanna Holsten

Key Points

  • To develop an automated framework for identifying module-sensitive components in life cycle inventory data.
  • Proposed a machine learning-based framework for classification.
  • Systematically combined substance-based information with textual similarity features.
  • Evaluated using over 13,000 components from the automotive industry.
  • Focused on module-sensitive components like printed circuit boards.
  • Achieved 97% classification accuracy using Random Forest for electronic components.
  • Demonstrated improved accuracy by combining substance-related information with name similarity metrics.
  • Showed that conventional approaches underestimated the environmental impacts of certain components.

Abstract

The increasing demand for data related to life cycle assessment, such as environmental product declarations, necessitates the development of more efficient approaches for the preparation and integration of life cycle assessment-relevant data. A central challenge in this context is data mapping, specifically the consistent linkage between foreground and background datasets. Of particular importance is the identification of so-called module-sensitive components, such as printed circuit boards, which consist of highly interdependent materials in terms of their environmental impact. The environmental impacts of these components are often underestimated by conventional approaches and typically require manual identification. To address this challenge, a machine learning-based framework is proposed to enable the automated classification of such components. To the best of the authors’ knowledge, this work presents the first framework that systematically combines substance-based information with textual similarity features for automated identification of module-sensitive components in life cycle inventory data preparation. The framework outlines how a sufficiently robust classification model can be prepared, optimized, and applied. It was evaluated using a dataset of more than 13,000 individual components from the automotive industry. The results demonstrate that combining substance-related information with name similarity metrics significantly improves classification accuracy. Among the tested algorithms, Random Forest achieved the best performance, reaching a classification accuracy of 97% for electronic components such as printed circuit boards, whereas Support Vector Machines and k-Nearest Neighbors yielded substantially lower accuracies. Overall, the proposed framework reduces manual effort, enhances data consistency, and supports the operationalization and scalability of life cycle inventory and, consequently, life cycle assessment preparation processes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sturm et al. (2026) studied this question.

synapsesocial.com/papers/69bb92df496e729e62980817https://doi.org/10.1016/j.jclepro.2026.147942
Ask AI
Helpful
Bookmark
Share
View Full Paper