PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 17, 2026Sakarya University Journal of Computer and Information Sciences0 citationsOpen Access

Performance Comparison of Multimodal Vision-Language Models in Classifying Turkish Dishes

View Full Paper
YBYunus Serhat Bıçakçı

Key Points

  • The research aims to benchmark multimodal vision-language models for classifying Turkish dishes using Turkish-language support.
  • Evaluated seven multimodal vision-language models with Turkish language support.
  • Used two image datasets of Turkish cuisine: TurkishFoods-15 and TurkishFoods-25.
  • Performed zero-shot evaluation without additional training or fine-tuning.
  • Utilized standardized Turkish prompts to assess model performance.
  • Aya Vision 32B achieved the highest weighted F1-score of 85.9% on TurkishFoods-15.
  • Gemma 3 27B led with a score of 76.7% on TurkishFoods-25.
  • Aya Vision 32B, Gemma 3 27B, Qwen2-VL 72B-AWQ, and InternVL3 38B were the most reliable models across all metrics.

Abstract

This study presents the first comprehensive benchmark of seven open-source multimodal vision-language models with Turkish language support—namely, Aya Vision 32B, Gemma 3 27B, InternVL3 38B, Qwen2-VL 72B-AWQ, Qwen2.5-VL 72B-AWQ, Cosmos-LLaVA, and Phi-4 Multimodal—on two image datasets of Turkish cuisine, TurkishFoods-15 and TurkishFoods-25. All models were evaluated zero-shot, without additional training or fine-tuning, utilizing a fully standardized Turkish system and user prompts. We report macro and weighted averages of accuracy, precision, recall, and F1-score, along with end-to-end inference time. Aya Vision 32B obtained the best weighted F1-score (85.9%) on TurkishFoods‑15, whereas Gemma 3 27B led on TurkishFoods‑25 (76.7%). Across metrics and datasets, Aya Vision 32B, Gemma 3 27B, Qwen2‑VL 72B‑AWQ, and InternVL3 38B formed the most reliable models. These results establish a solid reference for future work on culturally aware multimodal AI and demonstrate, for the first time, that vision-language models can categorize Turkish dishes without task‑specific training.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yunus Serhat Bıçakçı (2026) studied this question.

synapsesocial.com/papers/69b8f13ddeb47d591b8c6465https://doi.org/10.35377/saucis...1727583
Ask AI
Helpful
Bookmark
Share
View Full Paper