TrafficPerceiver: A multimodal large language model with reinforcement learning for unified challenging traffic scene perception
Senyun Kuang et al.
What the paper says
Understanding traffic scenes under diverse and challenging conditions is critical for intelligent transportation systems (ITS). Existing methods primarily focus on ideal scenarios and often lack the ability to perform fine-grained perception or respond to human instructions. To address these limitations, we propose TrafficPerceiver, a unified multimodal framework based on a multimodal large language model (MLLM) that jointly supports both image understanding and target-oriented segmentation. To enhance the model’s performance under adverse conditions such as rain, fog, and motion blur, we introduce a reinforcement learning optimization strategy based on group-relative policy optimization (GRPO), which encourages interpretable, instruction-following behavior. Additionally, we construct the challenging traffic scene understand ing (CTSU) dataset, a large-scale dataset tailored to challenging traffic environments, with dense annotations for both segmentation and instruction-response tasks. Extensive experiments on both the DRAMA-ROLISP and CTSU datasets demonstrate that TrafficPerceiver achieves state-of-the-art performance in both understanding and segmentation tasks.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.