PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 2026Future Internet0 citationsOpen Access

Aerial Image Analysis: When LLMs Assist (And When Not)

View Full Paper
SCSalvatore CalcagnoESErika ScalettaETEmiliano Tramontana

Key Points

  • The research aims to evaluate how effectively a large language model can identify and caption natural and man-made objects in aerial images.
  • Implemented tests using the Llama-4 LLM on a custom dataset of aerial images.
  • Assessed identification and captioning capabilities for tree categories, land, and roads.
  • Computed accuracy, precision, and recall metrics for evaluation.
  • Maverick, a variant of Llama-4, achieved a maximum accuracy of 58.6%.
  • The recall metric reached a maximum of 56.1%.
  • Findings highlight strengths in automation but indicate a need for significant improvements for reliable use.

Abstract

Large language models (LLMs) have shown remarkable results when tasked with the analysis and production of texts or images and for captioning images. Aerial images differ from other images since they exhibit many natural objects that have a highly variable color range and no clear contours. This paper reports to what extent an LLM, i.e., Llama-4, can be tasked with the identification and captioning in aerial images of natural objects, such as tree categories, uncultivated land, and some man-made objects, such as roads. This valuable automation is needed to scan large areas and detect the parts for which a sudden maintenance or an emergency intervention is due. Tests on the chosen LLM were performed against a custom image dataset built to overcome the limited availability of such a domain-specific aerial image set. To evaluate the identification and captioning results, the accuracy, precision and recall metrics were computed. The results given by a cutting-edge variant of Llama-4, namely Maverick, reveal its strengths and weaknesses in this context. Although it is remarkable that an out-of-the-box tool can give assistance in such a complex observation and detection task, substantial progress is needed for such a model to improve accuracy and constitute a reliable support, as accuracy is at most 58.6% and recall is at most 56.1%.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Calcagno et al. (2026) studied this question.

synapsesocial.com/papers/698434ebf1d9ada3c1fb3a52https://doi.org/10.3390/fi18020077
Ask AI
Helpful
Bookmark
Share
View Full Paper