PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 24, 2026IEEE Access0 citationsOpen Access

Mislabel Identification Using Transfer Learning-Based Ensemble Method

View Full Paper
MIMd Shariful IslamMKMin Jun KimPDPrashanta Dutta

Key Points

  • To develop a robust method for identifying mislabeled training data in machine learning applications.
  • Utilized transfer learning based on fine-tuned pretrained deep neural networks like ResNet and VGG.
  • Employed majority filtering and consensus filtering techniques for detection of mislabeled samples.
  • Validated the method on both the MNIST dataset and a nanopore dataset with injected mislabels.
  • Detected approximately 751 label inconsistencies in the MNIST dataset, consistent with existing estimates.
  • Recovered up to 100% of known corrupted labels during synthetic experiments.
  • Achieved superior accuracy compared to classical methods like KNN and k-means on balanced AAV datasets.

Abstract

Accurate labeling of training data is essential for reliable supervised machine learning, particularly in sensitive applications such as virus classification, autonomous driving, precision manufacturing, and medical diagnostics. However, the labeling process is labor-intensive and error-prone. Even widely used datasets such as MNIST and ImageNet contain numerous mislabeled samples. To address this challenge, we developed a transfer learning based ensemble method that identifies mislabeled data through majority filtering and consensus filtering using fine-tuned pretrained deep neural networks, including ResNet-50, ResNet-101, VGG-16, EfficientNet, MobileNet, and Inception. Our approach was first validated on the MNIST dataset, where the ensemble detected approximately 751 label inconsistencies, which closely aligns with previously reported estimates of mislabeled samples. Additional experiments with synthetically injected mislabels demonstrated that the method could recover up to 100% of known corrupted labels using majority and consensus voting strategies. The method was then applied to a highly pure adeno-associated virus (AAV) nanopore dataset, where artificial mislabels were introduced for evaluation; the ensemble successfully identified most mislabeled samples and correctly recovered their true labels. Experiments on balanced and unbalanced AAV datasets further showed improved performance on the balanced subset, where all injected mislabels were detected. Compared to classical filtering techniques such as KNN, k-means clustering, and advanced machine learning based mislabel detection (e.g., DivideMix), the proposed ensemble method demonstrated superior accuracy, stability, and true-label recovery, establishing it as a strong general-purpose mislabel detection framework—particularly well-suited for complex, fine-grained datasets such as nanopore signals and other biological measurement data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Islam et al. (2026) studied this question.

synapsesocial.com/papers/69eb07a4553a5433e34b31b5https://doi.org/10.1109/access.2026.3683309
Ask AI
Helpful
Bookmark
Share
View Full Paper