PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 10, 2026Agronomy Journal0 citations

Assessing the quality of crop variety data extraction from unstructured text sources using large language models

View Full Paper
MSM. A. SoldatkinaPTP. R. TsymbarovichDFD. S. Fomin

Key Points

  • This research aims to assess the effectiveness of large language models in extracting crop variety data from unstructured text sources.
  • Evaluation of large language models including GPT-3.5 Turbo, GPT-4o, and GPT-4o-mini.
  • Application of few-shot prompt-tuning and fine-tuning techniques.
  • Human assessment of extraction quality based on 250 variety descriptions.
  • GPT-4o model achieved the highest data extraction quality with an F1 score of 0.967.
  • GPT-4o recognized key parameters like morphological properties and abiotic stress tolerance effectively.
  • Text Coverage showed a strong correlation (0.405) with human assessments and identified a threshold for quality.

Abstract

Abstract A huge amount of agricultural information is on paper or electronic media in a non‐digitized and unstructured form. Obtaining detailed information about agricultural plants in a machine‐readable format is possible through extracting and analyzing data from unstructured text sources. This study evaluates the effectiveness of large language models (LLMs) in extracting crop variety data from Russian‐language agricultural registers, specifically the Russian State Register of Breeding Achievements (30,000+ varieties). GPT‐3.5 Turbo (where GPT is generative pretrained transformer), GPT‐4o, and GPT‐4o‐mini from OpenAI were selected as language models using few‐shot prompt‐tuning (FSPT) and fine‐tuning (FT) approaches. Human assessment of the data extraction quality was performed on 250 variety descriptions by three reviewers, who are specialists in the field of agriculture. The best result of data extraction was shown by the GPT‐4o model with FSPT (F1 = 0.967). The model led in recognizing the parameters “Morphological Properties,” “Yield and Production,” and “Abiotic Stress Tolerance,” but slightly lagged behind GPT‐4o‐mini with FT in the category “Tolerance to Diseases and Pests.” To scale the assessment of the quality of data extraction, a search for a relationship with the automatic metrics was carried out. Text Coverage demonstrated the strongest correlation with human assessment (cor = 0.405) and was set at 0.73 as the threshold for identifying unsatisfactory data extraction results. These results represent a significant advancement in automating the processing of agricultural data while maintaining high accuracy standards. This research contributes to the growing body of knowledge on applying LLMs to specialized agricultural domains and offers solution for large‐scale data extraction challenges.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Soldatkina et al. (2026) studied this question.

synapsesocial.com/papers/69af956970916d39fea4cf42https://doi.org/10.1002/agj2.70324
Ask AI
Helpful
Bookmark
Share
View Full Paper