This paper proposes a unified multimodal framework for automatic labeling and classification of tourist attractions by fusing images, texts, and geospatial cues with attention mechanisms. We encode attraction images with CNN backbones, descriptions with pretrained language models, and locations with a geospatial encoder, then project all modalities into a shared latent space and learn sample adaptive attention weights to emphasize the most informative signals. Based on this design, we develop a Multimodal Attention Fusion Network (MAFN) for efficient end-to-end prediction, and further introduce an Adaptive Multimodal Fusion Strategy (AMFS) that employs hierarchical attention and a regularization term to mitigate modality imbalance and improve robustness under noisy or missing inputs. Experiments on multiple tourist-attraction benchmarks demonstrate consistent improvements over strong baselines in accuracy, recall, F1, and AUC, while ablation studies confirm the effectiveness of both MAFN and AMFS. Quantitatively, our method achieves up to 90.34% accuracy, 89.78% recall, 89.23% F1, and 89.45% AUC across four benchmarks, and improves over the strongest baseline by up to +2.22 percentage points in accuracy and +2.23 percentage points in F1 (with consistent gains in recall and AUC).
Zhenhua Mei (Tue,) studied this question.