• A multimodal framework integrates ROI-focused vision models with ecological text descriptors • GPT-4o generates structured botanical descriptions enabling domain-adapted BERT training • Late-fusion of visual and textual posteriors significantly improves fine-grained classification • SAM2 and GroundingDINO enhance morphological signal by isolating plant inflorescences • Cross-platform mobile app operationalizes the classifier for real-time field identification Goldenrod ( Solidago spp.) is a genus of ecologically critical North American wildflowers that supports a diversity of pollinators, including bees, wasps, and and migratory insects such as Monarch butterflies. Yet species-level identification remains challenging due to extensive morphological similarity and overlapping distributions. Reliable automated classification would benefit ecological monitoring and biodiversity assessment, but visual models alone often struggle with fine-grained distinctions in natural imagery. This study introduces a multimodal framework that integrates image-based deep learning with text-derived ecological context to improve classification of 18 Solidago species. Images from the Global Biodiversity Information Facility (GBIF) are preprocessed using SAM2 and Grounding DINO to extract botanically relevant regions of interest. Five visual backbones such as VGG-19, ResNet-50, InceptionV3, Vision Transformer (ViT), and ConvNeXt are fine-tuned for species identification. In parallel, GPT-4o is used to generate structured botanical descriptions from GBIF metadata, which are then used to fine-tune a BERT-based classifier for textual inference. A late-fusion module integrates predictions from the visual and textual branches, leveraging complementary morphological and ecological cues. Across all architectures, multimodal fusion yields substantial improvements over image-only baselines, with ConvNeXt and ResNet achieving the highest overall accuracies (0.864 and 0.874) and strongest F1 scores (0.919 and 0.861). Extensive data analysis confirms that multimodal fusion consistently enhances discriminative power for morphologically similar taxa. To operationalize these findings, we also developed a cross-platform mobile application, enabling field-ready inference and real-time species exploration. Together, this work demonstrates that combining visual and contextual information provides a scalable and accurate pathway for identifying closely related plant species in natural environments.
Kazanjian et al. (Sun,) studied this question.