Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 5, 2025Open Access

Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability

View Full Paper
Ask AI
Bookmark
Share

Authors

NLNing LiJZJingran ZhangJCJ. Cui

Discussion

Loading...

Member takes

Overview

Empirical analysis shows GPT-4o's limitations in achieving semantic synthesis and instruction adherence.

Key Points

  • GPT-4o fails to consistently apply domain knowledge during image generation and editing tasks.
  • Evaluation metrics indicate persistent limitations, especially in conditional reasoning and instruction fidelity.
  • The systematic study across three dimensions reveals significant gaps in GPT-4o's multimodal generation capabilities.
  • Findings suggest a need for enhanced benchmarks and training strategies to improve context-aware generation.

Cite This Study

Li et al. (2025) studied this question.

synapsesocial.com/papers/68e24e59d6d66a53c2472eaehttps://doi.org/10.48550/arxiv.2504.08003
View Full Paper
Ask AI
Bookmark
Share