PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 24, 2026Applied Sciences0 citationsOpen Access

Mislabel Detection in Multi-Label Chest X-Rays via Prototype-Weighted Neighborhood Consistency in CoAtNet Embedding Space

View Full Paper
AGAriel GamboaMAMauricio ArayaCSClaudio Sotomayor

Key Points

  • This research aims to identify mislabeled annotations in multi-label chest X-rays using a training-free approach.
  • Utilized frozen CoAtNet features to create embeddings from chest X-rays.
  • Compared prototype voting weighted by distance against k-nearest neighbors methods.
  • Evaluated under various synthetic label-noise conditions.
  • LPV-DW-CS method achieved a maximum macro-AUROC of 0.8860.
  • kNN variants reached a Recall@budget of up to 99.44%.
  • Expert reviews revealed significant label inconsistencies, supporting the effectiveness of the neighborhood-consistency ranking.

Abstract

Large-scale chest X-ray (CXR) datasets often rely on report-derived or weak labels, introducing missing and incorrect annotations that can degrade downstream models and limit trust. We study training-free mislabel detection in multi-label CXRs by scoring neighborhood label consistency in a fixed embedding space. Using the NIH Chest X-ray Kaggle sample (5606 CXRs), we extract intermediate CoAtNet features and obtain 64-dimensional embeddings with a frozen CoAtNet backbone and a lightweight refinement head. On top of these embeddings, we compare kNN consistency baselines with distance weighting and label-set similarity against LPV-DW-CS, clustered prototype voting weighted by distance and cluster support. We evaluate three synthetic label-noise regimes with review budgets matched to the corruption rate: random single-label (5% and 20%), boundary-noise (20% corruption within the lowest-density 20% subset), and disjoint-label replacement (20% within that subset). LPV-DW-CS yields the highest downstream macro-AUROC after filtering top-ranked samples (up to 0.8860), while kNN variants achieve higher Recall@budget at the same review rates (up to 99.44%). An image-only expert Likert review of top-ranked real samples finds substantial label-set inconsistencies (54.1% for LPV-DW-CS-280-A; 60.5% for KNN-DW-LSS), supporting neighborhood-consistency ranking as a practical, training-free tool for targeted dataset auditing.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gamboa et al. (2026) studied this question.

synapsesocial.com/papers/69eb0b50553a5433e34b511fhttps://doi.org/10.3390/app16094067
Ask AI
Helpful
Bookmark
Share
View Full Paper