LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection
Lei Hao et al.
What the paper says
Effective deep feature extraction via feature-level fusion is crucial for multimodal object detection. However, previous studies often involve complex training processes that integrate modality-specific features by stacking multiple feature-level fusion units, leading to significant computational overhead. To address this issue, we propose a lightweight attention-guided self-modulation feature fusion network (LASFNet). The LASFNet adopts a single feature-level fusion unit to enable high-performance detection, thereby simplifying the training process. The attention-guided self-modulation feature fusion (ASFF) module in the model adaptively adjusts the responses of fused features at both global and local levels, promoting comprehensive and enriched feature generation. Additionally, a lightweight feature attention transformation module (FATM) is designed at the neck of LASFNet to enhance the focus on fused features and minimize information loss. Extensive experiments on three representative datasets demonstrate that our approach achieves a favorable efficiency-accuracy tradeoff. Compared to state-of-the-art methods, LASFNet reduced the number of parameters and computational cost by as much as 90% and 85%, respectively, while improving detection accuracy mean average precision (mAP) by 1%-3%. The code will be open-sourced at https://github.com/leileilei2000/LASFNet.
1 citation
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.16 × 0.4 = 0.06 |
| M · momentum | 0.53 × 0.15 = 0.08 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.