PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 3, 2026Safety0 citationsOpen Access

Risk Level Assessment and Impact Range Analysis of CCUS CO2 Pipeline Leakage Based on Machine Learning

View Full Paper
HZHaoyuan ZhangSWShuna WangXJXiaoping Jia

Key Points

  • The aim is to develop an integrated framework for assessing risk levels and impact distances of CO2 pipeline leakage using machine learning methods.
  • Constructed a scenario library with 4320 scenarios covering various conditions.
  • Implemented an integrated framework for threshold impact-distance calculation and risk-matrix mapping.
  • Used Extreme Gradient Boosting (XGBoost) for RiskLevel classification, compared with other machine learning models.
  • Applied a stratified 70%/30% train-test split for model training and testing.
  • XGBoost achieved a classification accuracy of 0.806 and a macro-F1 score of 0.825.
  • The recall for high-risk classes (RiskLevel 4-5) was 0.631.
  • Mean absolute errors (MAEs) for thresholds R1%, R4%, and R10% were 95, 62, and 41 m respectively, with R2 values between 0.795 and 0.814.
  • Analysis shows that impact distances increase with higher RiskLevels.

Abstract

In emergency decision-making for carbon capture, utilization, and storage (CCUS) CO2 pipeline leakage, risk levels and warning distances/impact ranges are often derived from different methodological systems—risk-matrix scoring versus mechanistic consequence modeling. Differences in threshold definitions and modeling assumptions make it difficult to align level assignment with distance boundaries for the same scenario, which in turn reduces the comparability and traceability of multi-scenario batch screening. To address this, this study proposes an integrated framework based on “threshold impact-distance calculation–risk-matrix mapping,” with physical consequence quantification as the main thread. A scenario library (N = 4320) covering phase state, leak aperture, operating conditions, and meteorological fields is constructed; impact distances corresponding to CO2 volume-fraction thresholds of 1%/4%/10% (R1%, R4%, R10%) are computed and then mapped to five RiskLevel classes under a unified rule set, enabling standardized synchronous outputs. The modeling tasks are formulated as RiskLevel classification and threshold-distance regression. Using a stratified 70%/30% train–test split, Extreme Gradient Boosting (XGBoost) is adopted as the primary model and compared with logistic regression (LR), support vector classification (SVC), ordinary least squares regression (OLS), and support vector regression (SVR). Results show that XGBoost achieves an accuracy of 0.806 and a macro-F1 of 0.825 for RiskLevel classification, with a recall of 0.631 for the high-risk classes (RiskLevel 4–5), and yields mean absolute errors (MAEs) of 95/62/41 m for R1%/R4%/R10% regression with coefficient of determination (R2) values of 0.795–0.814. Distributional analysis further indicates that threshold impact distances increase overall with higher RiskLevel, while dispersion becomes larger at higher levels. Accordingly, a parallel representation of “RiskLevel + multi-threshold rings” is recommended to support coordinated graded control and zoned warning delineation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69cf5dd55a333a821460be02https://doi.org/10.3390/safety12020044
Ask AI
Helpful
Bookmark
Share
View Full Paper