PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026Big Data and Cognitive Computing0 citationsOpen Access

Lithology Identification from Well Logs via Meta-Information Tensors and Quality-Aware Weighting

View Full Paper
WCWenxuan ChenGZGuoyun ZhongFDFan Diao

Key Points

  • This research aims to improve lithology identification from well logs despite challenges like missing data and class imbalance.
  • Developed a framework combining robust feature engineering with quality-aware XGBoost.
  • Used sentinel values and meta-information tensors to encode missing data patterns.
  • Implemented a sliding-window context to create auxiliary features.
  • Introduced a quality-aware sample-weighting strategy to mitigate training bias.
  • Improved weighted F1 score from 0.66 to 0.73 compared to baseline models.
  • Enhanced Boundary F1 score and geological penalty score.
  • Demonstrated advantages of handling data incompleteness explicitly over traditional cleaning methods.

Abstract

In practical well-logging datasets, severe missing values, anomalous disturbances, and highly imbalanced lithology classes are pervasive. To address these challenges, this study proposes a well-logging lithology identification framework that combines Robust Feature Engineering (RFE) with quality-aware XGBoost. Instead of relying on interpolation-based data cleaning, RFE uses sentinel values and a meta-information tensor to explicitly encode patterns of missingness and anomalies, and incorporates sliding-window context to transform data defects into discriminative auxiliary features. In parallel, a quality-aware sample-weighting strategy is introduced that jointly accounts for formation boundary locations and label confidence, thereby mitigating training bias induced by long-tailed class distributions. Experiments on the FORCE 2020 lithology prediction dataset demonstrate that, relative to baseline models, the proposed method improves the weighted F1 score from 0.66 to 0.73, while Boundary F1 and the geological penalty score are also consistently enhanced. These results indicate that, compared with traditional workflows that rely solely on data cleaning, explicit modeling of data incompleteness provides more pronounced advantages in terms of robustness and engineering applicability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2026) studied this question.

synapsesocial.com/papers/698433f6f1d9ada3c1fb18fchttps://doi.org/10.3390/bdcc10020047
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Focal Loss for Dense Object Detection2017 · 27,143 citations
  2. 2Extremely missing numerical data in Electronic Health Records for machine learning can be managed through simple imputation methods considering informative missingness: A comparative of solutions in a COVID-19 mortality case study2023 · 28 citations
  3. 3Evaluating missing data handling methods for developing building energy benchmarking models2024 · 17 citations
  4. 4Enhanced cross-domain lithology classification in imbalanced datasets using an unsupervised domain Adversarial Network2024 · 9 citations
  5. 5Cross‐validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure2016 · 3,000 citations