Multimodal Sentiment and Emotional Analysis in Short Video Dissemination
Rong He
What the paper says
Short videos dominate online content, yet sentiment's role in virality is understudied. This research proposes a multimodal framework integrating visual, textual, and interactional data from 127,843 anonymized videos. It extracts emotional features—visual (color saturation, brightness), textual sentiment (FinBERT, -1 to 1), and interactional feedback—using deep learning for fusion and structural equation modeling for causal analysis. Results show positive text sentiment with high brightness boosts shares by 18%; dwell time mediates emotional resonance and sharing. Emotional drivers vary by content: entertainment relies on visual stimuli, education on positive text sentiment, news on comment signals. The multimodal model achieves 89.7% accuracy in predicting dissemination, outperforming unimodal models by 9–13%. Findings enhance service information systems by optimizing recommendations, user experience, and public opinion management, offering actionable insights for aligning content with user emotions to improve engagement sustainability.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.