What question did this study set out to answer?

To develop a multimodal framework that enhances music auto-tagging by integrating various data sources and graph learning techniques.

April 12, 2026Open Access

A multimodal graph-based music auto-tagging framework: integrating social and content intelligence

Key Points

To develop a multimodal framework that enhances music auto-tagging by integrating various data sources and graph learning techniques.
Developed MuCoGraph, a hybrid learning framework for music auto-tagging.
Integrated lyrics, user comments, and audio spectrograms in the model.
Utilized three types of graph neural networks: preference, group, and tag co-occurrence graphs.
Applied hierarchical co-attention mechanisms to enhance cross-modal features.
MuCoGraph shows significant improvements over baseline methods in music auto-tagging.
Achieved 12% gain in F1-score, 11.6% in NDCG, and 11.8% in MAP compared to the best baseline.
The model maintains quality across larger tag lists, indicating robust performance in recommendations.

Abstract

Abstract With the rapid growth of music streaming platforms, effective music auto-tagging has become crucial for Music Information Retrieval (MIR) and recommendation. However, existing approaches face significant limitations: single-modality methods, which use only audio or text features, fail to capture the rich semantic diversity of music tags, while current multimodal approaches overlook critical interactive relationships beyond music content. Moreover, most studies ignore the co-occurrence dependencies among tags, which are essential for multi-label prediction. To address these challenges, we propose MuCoGraph, a novel multimodal graph-based hybrid learning framework for music auto-tagging. Our approach integrates multiple data modalities—lyrics, user comments, and audio spectrograms—with three heterogeneous graph neural networks: preference graphs that capture artist-listener interactions, group graphs that model content similarities, and tag co-occurrence graphs that learn label dependencies. The framework employs hierarchical co-attention mechanisms that enable cross-modal feature enhancement, allowing graph-based features to strengthen textual and audio representations through mutual learning. Experiments were conducted on a real-world dataset, which we integrated from multiple online platforms, demonstrating that MuCoGraph outperforms all the compared baseline methods in music auto-tagging. Notably, MuCoGraph achieves the most substantial improvements in top-12 recommendations, with relative gains of 12% in F1-score, 11.6% in NDCG, and 11.8% in MAP compared to the best baseline. This demonstrates progressively greater performance advantages as the recommendation scope increases, highlighting its enhanced ability to maintain quality across extended tag lists. These performance gains provide practical benefits for music platforms, including enhanced user engagement and more efficient content management processes. Furthermore, ablation studies demonstrate the critical contribution of each model component, particularly showing that graph-based features effectively improve both textual and audio representations through cross-modal enhancement.

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Authors

Yang Huang

Duen-Ren Liu

Yi‐Hsuan Chen

Journals

Information Technology and Management

Actions

Institutions

National Yang Ming Chiao Tung University

National Chung Hsing University

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

A multimodal graph-based music auto-tagging framework: integrating social and content intelligence

Key Points

Abstract

Citation Network

Connected Papers

Discussion

Authors

Journals

Actions

Institutions

References and Citations

Citation Network

Connected Papers

Discussion

Cite this study

Also consider