PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 24, 2026Iconic Research and Engineering Journals0 citations

AI-Driven Synesthetic Music Visualizer: Real-Time Cross-Modal Audio-to-Visual Translation Using Machine Learning

MSMohammed Yusoof SHNHarini C NHSHarini S

Key Points

  • The research aims to create a real-time music visualizer that translates audio signals into visual representations using AI.
  • Implementing signal processing techniques to extract audio features from MP3 or WAV tracks.
  • Training and benchmarking six machine learning models on a specialized audio-visual mapping dataset.
  • Deploying the visualizer as a browser-accessible application using GPU acceleration.
  • The Gradient Boosting model achieved the highest classification performance with an F1-score of 88.8%.
  • The system operates at an average inference latency of 29 ms, allowing for real-time synchronization.
  • The visual parameters generated include color palettes and shapes, rendered at 62 frames per second.

Abstract

This paper presents the design, implementation, and evaluation of an AI-Driven Synesthetic Music Visualizer — a real-time system that computationally emulates the neurological phenomenon of synesthesia by translating auditory signals into semantically congruent, dynamic visual art. Audio tracks in MP3 or WAV format are ingested and decomposed into perceptual acoustic features — including Root Mean Square (RMS) energy, spectral centroid, chroma vector, tempo, and spectral rolloff — through Short-Time Fourier Transform (STFT)-based signal processing using the Librosa library. Six machine learning architectures are trained and benchmarked on an annotated audio-visual mapping corpus: Multi-Layer Perceptron (MLP), Long Short-Term Memory (LSTM), Random Forest, Support Vector Machine (SVM), K-Nearest Neighbours (KNN), and Gradient Boosting. The Gradient Boosting model achieves the highest classification performance with an F1-score of 88.8 % and an average inference latency of 29 ms — well within the perceptual synchronisation budget. Predicted visual parameters (colour palette, shape morphology, animation velocity) are forwarded to a GPU-accelerated OpenGL rendering engine sustaining 62 frames per second on commodity hardware. The complete pipeline is deployed as a browser-accessible Gradio application. Results demonstrate that intelligent cross-modal synthesis is achievable in genuine real time, opening avenues for generative art, live performance, and assistive technology for hearing-impaired users.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

S et al. (2026) studied this question.

synapsesocial.com/papers/69eb08ef553a5433e34b3a03https://doi.org/10.64388/irev9i9-1715778
Ask AI
Helpful
Bookmark
Share
View Full Paper