VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction

Yancheng Ling et al.

Communications in Transportation Research2026https://doi.org/10.26599/commtr.2026.9640009article
ABDC C
Weight
0.50

What the paper says

Pedestrian crossing intention prediction is crucial for autonomous driving. While existing models have achieved high accuracy, their generalization and robustness remain limited, hindering their performance in real-world scenarios. To overcome these limitations, we introduce the LVLMPed-CoT, a large vision language model (LVLM) that incorporates a chain-of-thought (CoT) mechanism to  enhance pedestrian crossing intention prediction. It takes multimodal data as input and employs data distillation along with a two-stage fine-tuning strategy to elicit the implicit CoT capability of a lightweight vision-language model for enhanced perception, reasoning, an d prediction. The unified LVLMPed-CoT is trained on a joint open-source dataset (JAAD and PIE) and achieves superior or comparable performance to state-of-the-art models on both large-scale public datasets. The ablation study validates the contribution of the CoT prompt design and the two-stage fine-tuning strategy to the model's performance. Further analysis investigates the impact of input data sequence length and image quality on both accuracy and inference time, as well as the interpretability of the enhanced CoT reasoning ability achieved through fine-tuning.  

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.26599/commtr.2026.9640009

Or copy a formatted citation

@article{yancheng2026,
  title        = {{VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction}},
  author       = {Yancheng Ling et al.},
  journal      = {Communications in Transportation Research},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.26599/commtr.2026.9640009},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

VLMPed-CoT: A large vision-language model with a chain-of-thought mechanism for pedestrian crossing intention prediction

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.