PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 2026World Wide Web0 citationsOpen Access

Cluster-SCP: Similarity and Contrastive Learning to Enhance Pseudo Labels for Fine Tuning under Few Labels

AAAbdullah AlsuhaibaniAAAbdulrahman AlalawiIRImran Razzak

Key Points

  • The aim is to enhance pseudo labeling for fine tuning language models with limited labeled data.
  • Developed Cluster-SCP framework for improved pseudo labeling.
  • Utilized K-Means for clustering initialization.
  • Implemented iterative refinement with Intermediate Pseudo Labels.
  • Executed embedding reassignment and contrastive learning for cluster merging.
  • Achieved at least a 5.3% accuracy improvement in fine tuning.
  • Reduced the number of clusters by 22% while maintaining performance.
  • Cluster-SCP outperformed baseline and state-of-the-art methods on benchmark datasets.

Abstract

Language models often underperform when fine tuned with limited labelled data. To reduce dependence on costly annotation, recent studies have explored clustering and pseudo labelling to leverage unlabeled text with only a few labelled examples for fine tuning. However, these methods face persistent challenges: clustering errors and pseudo-label mismatches can degrade fine-tuning performance. We propose Cluster-SCP, a novel framework that improves intra-cluster coherence and reduces cluster quantity to generate more reliable pseudo labels. The framework begins with K-Means initialization and applies an iterative refinement process, Intermediate Pseudo Labels, comprising two stages: embedding reassignment to minimize clustering errors, and cluster merging via contrastive learning with graph readouts. The refined pseudo labels are first used to fine tune language model, which is subsequently refined using a small set of labels. Experiments on three benchmark datasets show that Cluster-SCP consistently outperforms baseline and state-of-the-art methods, achieving at least a 5.3% accuracy improvement while reducing the number of clusters by 22%.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alsuhaibani et al. (2026) studied this question.

synapsesocial.com/papers/69e320e740886becb6540101https://doi.org/10.1007/s11280-026-01415-w
Ask AI
Helpful
Bookmark
Share
View Full Paper