PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs

View Full Paper
YLYuanshuai LiYYYuping YanJTJunfeng Tang

Key Points

  • SCPO reduces visual hallucinations by up to 62.9% while enhancing model factuality and capabilities.
  • The framework utilizes a progressive curriculum derived from a specialized dataset focusing on semantic differences.
  • Performance assessments on LLaVA models reveal SCPO outperforms traditional methods across multiple hallucination benchmarks.
  • The approach integrates symmetry and a dynamic reference model to improve the learning process in MLLM alignment.

Abstract

Multimodal Large Language Models (MLLMs) have significantly improved the performance of various tasks, but continue to suffer from visual hallucinations, a critical issue where generated responses contradict visual evidence. While Direct Preference Optimization(DPO) is widely used for alignment, its application to MLLMs often fails to capture fine-grained semantic differences and encourages shortcut learning. To address these challenges, we propose Semantic Curriculum Preference Optimization (SCPO), a novel framework for MLLM alignment. SCPO employs a progressive, easy-to-hard curriculum built upon our Semantic Curriculum Preference Pairs dataset, which provides fine-grained semantic contrasts sorted by difficulty. This curriculum is trained with a dynamic reference model and a novel symmetric, bidirectional objective to facilitate simultaneous learning from both textual and visual preferences. To our knowledge, SCPO is the first framework to unify semantics, symmetry, and curriculum for MLLMs alignment, effectively mitigating visual hallucinations. Extensive experiments on LLaVA models across various scales and versions validate that SCPO demonstrates superior performance compared to baseline models on multiple hallucination benchmarks, reducing the hallucination rate by up to 62.9%. Moreover, evaluations on generalized benchmarks show that SCPO improves factuality while preserving general capabilities, with its performance remaining stable across general vision-language benchmarks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcce8d54a28a75cf1c1dhttps://doi.org/10.48550/arxiv.2509.24491
Ask AI
Helpful
Bookmark
Share
View Full Paper