PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 2026Transactions of the Association for Computational Linguistics0 citationsOpen Access

Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023

View Full Paper
THTim HsuYHYi-Li HsuSRShaurya Rohatgi

Key Points

  • This research explores the effectiveness of large multimodal models for generating captions for scientific figures.
  • Overview of the first SciCap Challenge
  • Evaluation of model performances on the SciCap dataset
  • Analysis of preferences by professional editors
  • Professional editors preferred GPT-4V captions over all other models
  • Models showed varying performance in caption generation across diverse figure types

Abstract

Abstract Since the SciCap dataset’s launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the first SciCap Challenge took place, inviting global teams to use an expanded SciCap dataset to develop models for captioning diverse figure types across various academic fields. At the same time, text generation models advanced quickly, with many powerful pre-trained large multimodal models (LMMs) emerging that showed impressive capabilities in various vision-and-language tasks. This paper presents an overview of the first SciCap Challenge and details the performance of various models on its data, capturing a snapshot of the field’s state. We found that professional editors overwhelmingly preferred figure captions generated by GPT-4V over those from all other models and even the original captions written by authors. Following this key finding, we conducted detailed analyses to answer this question: Have advanced LMMs solved the task of generating captions for scientific figures?

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hsu et al. (2026) studied this question.

synapsesocial.com/papers/69b4b9fb18185d8a398023b8https://doi.org/10.1162/tacl.a.653
Ask AI
Helpful
Bookmark
Share
View Full Paper