KEPT: Knowledge-enhanced prediction of trajectories from consecutive driving frames with vision-language models

Yujin Wang et al.

Communications in Transportation Research2026https://doi.org/10.26599/commtr.2026.9640012article
ABDC C
Weight
0.50

What the paper says

Accurate short-horizon trajectory prediction is crucial for safe and reliable autonomous driving. However, existing vision-language models (VLMs) often fail to accurately understand driving scenes and generate trustworthy trajectories. To address this challenge, this paper introduces KEPT, a knowledge-enhanced VLM framework that predicts ego trajectories directly from consecutive front-view driving frames. KEPT integrates a temporal frequency–spatial fusion (TFSF) video encoder, which is trained via self-supervised learning with hard-negative mining, with a k-means & HNSW retrieval-augmented generation (RAG) pipeline. Retrieved prior knowledge is added into chain-of-thought (CoT) prompts with explicit planning constraints, while a triple-stage fine-tuning paradigm aligns the VLM backbone to enhance spatial perception and trajectory prediction capabilities. Evaluated on nuScenes dataset, KEPT achieves the best open-loop performance compared with baseline methods. Ablation studies on fine-tuning stages, Top-K value of RAG, different retrieval strategies, vision encoders, and VLM backbones are conducted to demonstrate the effectiveness of KEPT. These results indicate that KEPT offers a promising, data-efficient way toward trustworthy trajectory prediction in autonomous driving. 

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.26599/commtr.2026.9640012

Or copy a formatted citation

@article{yujin2026,
  title        = {{KEPT: Knowledge-enhanced prediction of trajectories from consecutive driving frames with vision-language models}},
  author       = {Yujin Wang et al.},
  journal      = {Communications in Transportation Research},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.26599/commtr.2026.9640012},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

KEPT: Knowledge-enhanced prediction of trajectories from consecutive driving frames with vision-language models

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.