PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 22, 2026Journal of Information & Knowledge Management0 citations

Improved Le-SqueezeNet Model for Fine-Grained Multimodal Sentiment Analysis with Attention Mechanism-Based Aspect Extraction

View Full Paper
PPPradipta PatilDr. D. Y. Patil Medical College, Hospital and Research CentreSGSunil GuptaJaipur National University

Key Points

  • The aim is to improve sentiment analysis accuracy by integrating text, image, and audio inputs using a new model.
  • Inputs include text, images, and audio processed through tokenization, Gaussian filtering, and low-pass filtering.
  • Key features are extracted using DL-ATE, TF-IDF for text, PHOG and LGIP for images, and EMD for audio.
  • Features from all modalities are integrated using a Serial-based Maximum Information Feature Fusion approach.
  • Sentiment predictions are refined using advanced deep learning models like PAM-LNet and SqueezeNet.
  • The PAM-LNet-MSA achieved an accuracy rate of 0.956.
  • The model demonstrated a precision rate of 0.936.
  • It also reported an F-measure of 0.935, showing significant improvement over traditional methods.

Abstract

Conventional sentiment analysis focuses on text-level mining, often leading to lower accuracy. It uses computational linguistics and Natural Language Processing (NLP) to recognise and analyse emotions. As a result, there is growing interest in speech and facial expression recognition to improve accuracy, as single-modal analysis no longer meets modern needs. Therefore, integrating multiple modalities is essential for capturing richer emotional cues and achieving more reliable sentiment analysis. This paper proposes a novel Position Attention Module-assisted LeNet-based Multimodal Sentiment Analysis (PAM-LNet-MSA) framework. The process begins with the input phase, where data from text, images, and audio are processed. For text, the process involves tokenisation and stemming. The images are filtered using a Gaussian filter to remove noise, while the audio is cleaned using a low-pass filter to eliminate high-frequency noise. The next phase is feature extraction, where the system extracts key features from each modality. For text, Deep Learning-based Aspect Term Extraction (DL-ATE) and Term Frequency-Inverse Document Frequency (TF-IDF) are extracted. Images are analysed using Pyramid Histogram of Oriented Gradients (PHOG) and Local Gabor Increasing Pattern (LGIP). For audio, Empirical Mode Decomposition (EMD) and spectral features capture the frequency patterns of the sound. After extracting features, the system integrates the features from all three modalities using a Serial-based Maximum Information Feature Fusion (S-MIFF) approach. This unified set of features is then passed to a sentiment analysis model, where advanced deep learning architectures like Position Attention Module-assisted LeNet (PAM-LNet) and SqueezeNet are employed to refine sentiment predictions. The PAM-LNet-MSA significantly outperforms the traditional strategies with a greater accuracy rate of 0.956, precision of 0.936 and Formula: see text-measure of 0.935.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Patil et al. (2026) studied this question.

synapsesocial.com/papers/699a9d3c482488d673cd3017https://doi.org/10.1142/s0219649226500048
Ask AI
Helpful
Bookmark
Share
View Full Paper