PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 28, 2026IEEE Transactions on Neural Systems and Rehabilitation Engineering0 citationsOpen Access

Edge-Based Vision-Language Assistive System for the Visually Impaired: A Quantized VLM Approach

View Full Paper
BABatyr ArystanbekovHVHüseyin Atakan VarolAYAdnan Yazici

Key Points

  • The aim is to develop a self-contained, wearable assistive system for real-time scene interpretation for visually impaired individuals.
  • Integration of a quantized vision-language model with speech recognition and synthesis.
  • Utilization of NVIDIA Jetson Orin NX and Raspberry Pi for on-device processing.
  • Evaluation with 28 participants comparing the system to a baseline image captioning model.
  • Achieved a 25% increase in image identification accuracy compared to the baseline model.
  • Demonstrated minimal accuracy degradation (2.6% and 0.9%) in large-scale benchmarks.
  • Maintained practical latency of 4-5 seconds per query, indicating usability.

Abstract

This paper presents a novel edge-deployable assistive system for visually impaired individuals, powered by vision-language models (VLMs). Existing cloud-based solutions for image captioning suffer from latency, dependency on internet connectivity, and overly simplistic scene descriptions that fail to convey the rich contextual information needed for meaningful real-world navigation. To address these challenges, we propose a self-contained, wearable system that performs real-time scene interpretation on-device without cloud reliance. The system integrates a quantized version of the LLaVA-NeXT-13B VLM with speech recognition (Whisper) and speech synthesis (PIPER), running on an NVIDIA Jetson Orin NX and Raspberry Pi-based input module. Our framework emphasizes intuitive user interaction through a button-based interface and Bluetooth audio output, minimizing cognitive load. We validate the quantization approach through large-scale benchmarks, such as VizWiz-VQA and VQAv2, demonstrating minimal accuracy degradation (2.6% and 0.9%, respectively) compared to the original model. User evaluations involving 28 participants compared the system to a baseline image captioning model. Objective results demonstrated a 25% increase in image identification accuracy. The system achieved high usability scores and maintained practical latency (4-5 seconds per query), supporting real-world feasibility. This work advances the development of scalable, interpretable, and accessible AI-driven assistive technologies, enabling greater independence and interaction for the visually impaired.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Arystanbekov et al. (2026) studied this question.

synapsesocial.com/papers/6a17daf83fad632b0f9d7cd5https://doi.org/10.1109/tnsre.2026.3696447
Ask AI
Helpful
Bookmark
Share
View Full Paper