PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 6, 2026ACM Transactions on Internet of Things0 citations

Multi-Perspective Visual Contrastive Decoding for Reliable Assistance

View Full Paper
BPBocheng PanHSHailong ShiXGXingyu Gao

Key Points

  • The study aims to improve the effectiveness of MLLMs for individuals with blindness and low vision by addressing image processing challenges.
  • Developed a novel framework called MPVCD for visual contrastive decoding.
  • Implemented three perspectives: Noise Contrastive Decoding, Retrieval Contrastive Decoding, and Focus Contrastive Decoding.
  • Optimized predictions through Adaptive Perspective Integration for better token selection.
  • Demonstrated reduced hallucinations in generated visual descriptions across varied datasets.
  • Achieved more accurate and reliable visual descriptions for assisting BLV users.

Abstract

Multimodal Large Language Models (MLLMs) offer promising capabilities for assisting individuals with blindness and low vision (BLV), but their effectiveness is compromised when processing BLV-captured images, which typically suffer from three fundamental challenges: quality degradation, object incompleteness, and spatial misalignment. This paper presents MPVCD (Multi-Perspective Visual Contrastive Decoding), a novel framework that addresses these challenges through visual contrastive decoding techniques. MPVCD implements three specialized perspectives: Noise Contrastive Decoding addresses quality issues by comparing predictions between original and noise-injected images; Retrieval Contrastive Decoding tackles object incompleteness by retrieving semantically similar images from a memory bank; and Focus Contrastive Decoding resolves spatial misalignment by focusing on detected object regions. These perspectives are dynamically balanced through an Adaptive Perspective Integration that optimizes token selection based on prediction confidence. Our comprehensive experiments across diverse datasets demonstrate MPVCD’s effectiveness in reducing hallucinations under varied scenarios. By generating more accurate and reliable visual descriptions, MPVCD represents a significant advancement toward assistive technologies that BLV users can confidently rely on for environmental understanding and decision-making.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pan et al. (2026) studied this question.

synapsesocial.com/papers/698585678f7c464f23008ab2https://doi.org/10.1145/3785360
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection2023 · 49 citations
  2. 2VizWiz2010 · 594 citations
  3. 3Visual Hallucinations of Multi-modal Large Language Models2024 · 28 citations
  4. 4Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering2018 · 5,156 citations
  5. 5Women Also Snowboard: Overcoming Bias in Captioning Models2018 · 403 citations