Adaptive Temporal Expert Routing with Hierarchical Wavelet Enhancement for Multi-Modal Sequential Recommendation
Shiyu Liu et al.
What the paper says
Sequential recommendation systems have become essential for personalized services in e-commerce and content platforms. While recent research has extended these systems with multi-modal features, existing approaches face three major challenges. First, they inadequately model fine-grained temporal interval distributions, failing to discriminate between high-frequency short intervals and low-frequency long intervals. Second, uniform fusion in the time domain leads to semantic misalignment across modalities because it ignores their inherent differences in the frequency domain. Third, rigid fusion strategies without self-supervised constraints lead to limited representation quality and semantic drift from pre-trained embeddings. To address these issues, we propose ATHWE, an A daptive T emporal Expert Routing with H ierarchical W avelet E nhancement framework. ATHWE employs exponential saturation time mapping to generate temporally adaptive embeddings. These embeddings guide a sparse mixture of experts to model multi-scale user behavior dynamics. A hierarchical wavelet decomposition with band-specific gating selectively fuses complementary frequency components across modalities. Furthermore, contrastive learning and cluster-preserving objectives preserve semantic information during multi-modal fusion. Extensive experiments on multiple datasets validate the effectiveness of our framework. Our code is available at https://github.com/lulusiyuyu/ATHWE .
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.