LLM-Guided Multimodal Information Fusion With Hierarchical Spatio-Temporal Graph Network for Sentiment Analysis

Yujie Jin et al.

International Journal of Information Systems in the Service Sector2025https://doi.org/10.4018/ijisss.388002article
AJG 1
Weight
0.50

What the paper says

Multimodal sentiment analysis aims to attain a precise comprehension of emotions by integrating complementary textual, visual, and audio information. However, issues such as sentiment discrepancies between modalities, ineffective integration of multi-modal information, and the intricacy of order dependency significantly constrain the models' efficacy. The authors propose an LLM-guided Hierarchical Spatio-Temporal Graph Network (L-HSTGN). By multimodal large model feature enhancement, bidirectional spatio-temporal joint modeling, and dynamic gate fusion mechanism, they effectively address the aforementioned problems. Firstly, they produce cross-modal emotion pseudo-labels based on the multimodal large model, and the single-modal representation was optimized by combining adversarial regularization. Secondly, they develop a bidirectional spatio-temporal convolution module to concurrently extract local-global temporal characteristics and dynamic spatial correlations.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.4018/ijisss.388002

Or copy a formatted citation

@article{yujie2025,
  title        = {{LLM-Guided Multimodal Information Fusion With Hierarchical Spatio-Temporal Graph Network for Sentiment Analysis}},
  author       = {Yujie Jin et al.},
  journal      = {International Journal of Information Systems in the Service Sector},
  year         = {2025},
  doi          = {https://doi.org/https://doi.org/10.4018/ijisss.388002},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

LLM-Guided Multimodal Information Fusion With Hierarchical Spatio-Temporal Graph Network for Sentiment Analysis

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.