PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 30, 2026MethodsX0 citationsOpen Access

KeyCap3D: Keyword-Guided 3D Medical Image Captioning with Cross-Attention

View Full Paper
SSS. SupriyantoMSMuhammad Ibadurrahman Arrasyid SupriyantoHHHaviluddin Haviluddin

Key Points

  • This research aims to develop an automated radiological report generation system using 3D FLAIR MRI images.
  • Developed a keyword-guided cross-attention framework for image captioning
  • Integrated a transformer-based decoder for autoregressive caption generation
  • Utilized hierarchical keyword extraction with KeyBERT and BioBERT to enrich image representations
  • Trained on the BraTS2020 dataset with 369 glioma patients using NVIDIA RTX 3050 GPU
  • Achieved loss reduction from 4.16 to 1.33 during training
  • Evaluated models demonstrated BLEU-1 of 0.5359, BLEU-2 of 0.3969, and ROUGE-L of 0.5051
  • Generated captions accurately captured clinical information for decision support applications

Abstract

This study presents a keyword-guided cross-attention framework for automated radiological report generation from 3D FLAIR MRI brain tumor images. The architecture integrates M3D-CLIP image encoder with hierarchical keyword extraction using fine-tuned KeyBERT and BioBERT semantic embeddings in a 768-dimensional space. Six cross-attention layers fuse visual features with clinical keywords extracted across four hierarchical levels: abnormality type, lesion characteristics, anatomical location, and lateralization. A four-layer transformer decoder generates captions autoregressively. The BraTS2020 dataset containing 369 glioma patients paired with TextBraTS radiological descriptions was preprocessed with center-focused slice selection of 32 from 155 slices and spatial interpolation to 256 × 256 resolution. Training on NVIDIA RTX 3050 GPU for 15 epochs using AdamW optimizer achieved loss reduction from 4.16 to 1.33. Evaluation on 20 test samples demonstrated BLEU-1 of 0.5359, BLEU-2 of 0.3969, and ROUGE-L of 0.5051, with generated captions accurately capturing clinical information for decision support applications. • Multi-modal fusion through keyword-guided cross-attention integrating visual MRI features with hierarchical clinical terminology • Transformer-based autoregressive generation conditioned on enriched image-keyword representations • Comprehensive evaluation using BLEU and ROUGE metrics on brain tumor caption generation task

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Supriyanto et al. (2026) studied this question.

synapsesocial.com/papers/69c9c5a4f8fdd13afe0bda5chttps://doi.org/10.1016/j.mex.2026.103890
Ask AI
Helpful
Bookmark
Share
View Full Paper