Language-Aware Deep Learning for Crash Severity Prediction Modeling in Khmer Traffic Reports

Chhen Ey Khim et al.

Transportation Research Record2026https://doi.org/10.1177/03611981261430737article
ABDC B
Weight
0.50

What the paper says

Traffic crashes are a leading cause of death in low- and middle-income countries, where weak infrastructure and limited data hinder effective responses. Predicting crash injury severity is vital for emergency planning and policy; however, most machine learning models rely on English-language data, limiting their use in multilingual, low-resource settings. This is especially problematic for Khmer, Cambodia’s official language, which lacks word boundaries, has complex morphology, and suffers from scarce natural language processing resources. Standard models fail because of poor tokenization, semantic drift, and lack of script-specific representations. To address this, a Khmer-aware deep learning framework is proposed that integrates conditional random field-based tokenization, multigranular embeddings (character, subword, word), a dilated bidirectional long short-term memory with self-attention, and noise-robust classification to manage linguistic complexity and data variability. A labeled dataset of 1,074 Khmer-language traffic reports collected from eight Cambodian news outlets (2015–2024) is also introduced. The model achieves 95.2% accuracy, 0.952 precision, 0.952 recall, and 0.951 macro-F1, outperforming the best traditional model (eXtreme Gradient Boosting: 88.0% accuracy, 0.80 macro-F1) with nearly 60% lower error rate. Results confirm that language-specific design is essential for reliable severity prediction in low-resource languages. Exploratory analysis of media-reported crashes reveals that 40.7% were classified as fatal, 52.1% of fatalities occurred on national roads, and 73.6% involved motorcycle patterns reflective of reporting intensity rather than population-level risk. This work provides a reproducible pipeline to transform vernacular text into public health intelligence. By combining linguistic expertise with deep learning, it is demonstrated that inclusive, language-aware AI can turn local narratives into actionable, life-saving insights, setting a precedent for equitable road safety research in underserved regions.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1177/03611981261430737

Or copy a formatted citation

@article{chhen2026,
  title        = {{Language-Aware Deep Learning for Crash Severity Prediction Modeling in Khmer Traffic Reports}},
  author       = {Chhen Ey Khim et al.},
  journal      = {Transportation Research Record},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1177/03611981261430737},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Language-Aware Deep Learning for Crash Severity Prediction Modeling in Khmer Traffic Reports

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.