Multimodal Large Language Models (MLLMs) offer promising capabilities for assisting individuals with blindness and low vision (BLV), but their effectiveness is compromised when processing BLV-captured images, which typically suffer from three fundamental challenges: quality degradation, object incompleteness, and spatial misalignment. This paper presents MPVCD (Multi-Perspective Visual Contrastive Decoding), a novel framework that addresses these challenges through visual contrastive decoding techniques. MPVCD implements three specialized perspectives: Noise Contrastive Decoding addresses quality issues by comparing predictions between original and noise-injected images; Retrieval Contrastive Decoding tackles object incompleteness by retrieving semantically similar images from a memory bank; and Focus Contrastive Decoding resolves spatial misalignment by focusing on detected object regions. These perspectives are dynamically balanced through an Adaptive Perspective Integration that optimizes token selection based on prediction confidence. Our comprehensive experiments across diverse datasets demonstrate MPVCD’s effectiveness in reducing hallucinations under varied scenarios. By generating more accurate and reliable visual descriptions, MPVCD represents a significant advancement toward assistive technologies that BLV users can confidently rely on for environmental understanding and decision-making.
Pan et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: