When Multimodal Interactions Impair Prediction: A Novel Regularized Deep Learning Strategy

Gang Chen et al.

INFORMS Journal on Computing2026https://doi.org/10.1287/ijoc.2024.0794article
AJG 3ABDC A
Weight
0.50

What the paper says

Multimodal data are proliferating and hence flourishing data-driven business decision making, exemplified by short video attractiveness prediction (SVAP), multimodal review sentiment classification (MRSC), and multimodal data-based default risk prediction (DRP). However, when data of various modalities (e.g., text, graph, image, and video) are used jointly, they may mutually interact, adversely affecting prediction performance. To unravel and resolve the opaque conflicts in multimodal data, we formally conceptualize multimodal interactions and provide analytical insights for mitigating negative interactions at the feature, modality, and modality-wise instance levels. To better realize the predictive power of multimodal data, we propose a novel deep learning strategy named NIRMD (for negative interaction-regularized multimodal deep learning), which allows positive (negative) multimodal interactions to be effectively encouraged (mitigated) in a learnable nonlinear representation space. Empirical evaluation in three case studies involving SVAP, MRSC, and DRP, respectively, shows that the prediction performance of state-of-the-art multimodal deep learning methods can be enhanced by incorporating NIRMD. Exploratory (i.e., ablation, feature contribution, and case) analyses render evidence of NIRMD’s effectiveness in mitigating negative multimodal interactions. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: G. Chen was supported by the National Natural Science Foundation of China [Grants 72522010, 72301239, and 72394371]. S. Xiao was supported by the National Natural Science Foundation of China [Grants 72301194, 72495133, and 72472058]. C. Zhang was supported by the National Natural Science Foundation of China [Grants 72271059 and 72571071]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2024.0794 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2024.0794 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1287/ijoc.2024.0794

Or copy a formatted citation

@article{gang2026,
  title        = {{When Multimodal Interactions Impair Prediction: A Novel Regularized Deep Learning Strategy}},
  author       = {Gang Chen et al.},
  journal      = {INFORMS Journal on Computing},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1287/ijoc.2024.0794},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

When Multimodal Interactions Impair Prediction: A Novel Regularized Deep Learning Strategy

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.