PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 24, 2026Discover Artificial Intelligence0 citationsOpen Access

Efficient multimodal learning using BERT and vision transformers for visual question answering on peripheral blood cells

FSFaheem ShehzadCMCiro MennellaACAndrea Calimera

Key Points

  • The study aims to develop an efficient multimodal learning framework for visual question answering related to peripheral blood cells.
  • Developed a dual-stream architecture combining BERT and vision transformers.
  • Constructed a dataset with expert-annotated question-answer pairs for peripheral blood cell images.
  • Conducted extensive experiments to evaluate performance against state-of-the-art models.
  • Achieved a WUPS score of 0.95.
  • Obtained an F1 score of 0.94.
  • Reached an accuracy of 0.95, surpassing previous multimodal models.

Abstract

Peripheral blood cells are routinely examined in clinical practice, and answering clinically relevant questions about these cells is essential for decision support. In this work, we present an efficient multimodal learning framework for visual question answering on peripheral blood cells, combining BERT for textual representation with vision transformers for visual feature extraction. The proposed dual-stream architecture enables effective fusion of linguistic and visual information tailored to hematology data. To support this task, we construct a dedicated dataset by enriching an existing peripheral blood cell image collection with expert-annotated question–answer pairs. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art multimodal models, including BLIP and ViLT, achieving a WUPS score of 0.95, an F1 score of 0.94, and an accuracy of 0.95. These results establish a new benchmark for visual question answering in hematology and highlight the potential of efficient transformer-based multimodal models to enhance automated blood cell analysis and clinical interpretation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shehzad et al. (2026) studied this question.

synapsesocial.com/papers/699d405ade8e28729cf6557bhttps://doi.org/10.1007/s44163-026-01011-x
Ask AI
Helpful
Bookmark
Share
View Full Paper