TrafficPerceiver: A multimodal large language model with reinforcement learning for unified challenging traffic scene perception

Senyun Kuang et al.

Communications in Transportation Research2026https://doi.org/10.26599/commtr.2026.9640008article
ABDC C
Weight
0.50

What the paper says

Understanding traffic scenes under diverse and challenging conditions is critical for intelligent transportation systems (ITS). Existing methods primarily focus on ideal scenarios and often lack the ability to perform fine-grained perception or respond to human instructions. To address these limitations, we propose TrafficPerceiver, a unified multimodal framework based on a multimodal large language model (MLLM) that jointly supports both image understanding and target-oriented segmentation. To enhance the model’s performance under adverse conditions such as rain, fog, and motion blur, we introduce a reinforcement learning optimization strategy based on group-relative policy optimization (GRPO), which encourages interpretable, instruction-following behavior. Additionally, we construct the challenging traffic scene understand ing (CTSU) dataset, a large-scale dataset tailored to challenging traffic environments, with dense annotations for both segmentation and instruction-response tasks. Extensive experiments on both the DRAMA-ROLISP and CTSU datasets demonstrate that TrafficPerceiver achieves state-of-the-art performance in both understanding and segmentation tasks.  

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.26599/commtr.2026.9640008

Or copy a formatted citation

@article{senyun2026,
  title        = {{TrafficPerceiver: A multimodal large language model with reinforcement learning for unified challenging traffic scene perception}},
  author       = {Senyun Kuang et al.},
  journal      = {Communications in Transportation Research},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.26599/commtr.2026.9640008},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

TrafficPerceiver: A multimodal large language model with reinforcement learning for unified challenging traffic scene perception

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.