Improved Le-SqueezeNet Model for Fine-Grained Multimodal Sentiment Analysis with Attention Mechanism-Based Aspect Extraction
Pradipta Patil & Sunil Gupta
What the paper says
Conventional sentiment analysis focuses on text-level mining, often leading to lower accuracy. It uses computational linguistics and Natural Language Processing (NLP) to recognise and analyse emotions. As a result, there is growing interest in speech and facial expression recognition to improve accuracy, as single-modal analysis no longer meets modern needs. Therefore, integrating multiple modalities is essential for capturing richer emotional cues and achieving more reliable sentiment analysis. This paper proposes a novel Position Attention Module-assisted LeNet-based Multimodal Sentiment Analysis (PAM-LNet-MSA) framework. The process begins with the input phase, where data from text, images, and audio are processed. For text, the process involves tokenisation and stemming. The images are filtered using a Gaussian filter to remove noise, while the audio is cleaned using a low-pass filter to eliminate high-frequency noise. The next phase is feature extraction, where the system extracts key features from each modality. For text, Deep Learning-based Aspect Term Extraction (DL-ATE) and Term Frequency-Inverse Document Frequency (TF-IDF) are extracted. Images are analysed using Pyramid Histogram of Oriented Gradients (PHOG) and Local Gabor Increasing Pattern (LGIP). For audio, Empirical Mode Decomposition (EMD) and spectral features capture the frequency patterns of the sound. After extracting features, the system integrates the features from all three modalities using a Serial-based Maximum Information Feature Fusion (S-MIFF) approach. This unified set of features is then passed to a sentiment analysis model, where advanced deep learning architectures like Position Attention Module-assisted LeNet (PAM-LNet) and SqueezeNet are employed to refine sentiment predictions. The PAM-LNet-MSA significantly outperforms the traditional strategies with a greater accuracy rate of 0.956, precision of 0.936 and [Formula: see text]-measure of 0.935.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.