Performance unfairness of large language models in cross-language fact-checking

Dandan Wang et al.

Information Processing and Management2026https://doi.org/10.1016/j.ipm.2026.104616article
AJG 2
Weight
0.37

What the paper says

• Scalable computable evaluation system for LLM performance in fact-checking. • Inequality quantification across languages. • Checking-worthiness scoring and checking-authenticity verification. • Prompts in different types of language combination with different effects. • Role-restricted prompt engineering and model fine-tuning alleviate unfairness. Large language models (LLMs) are increasingly used for automated fact-checking, yet their performance often varies across languages, raising global fairness concerns. This study evaluated cross-language inequality in LLM-based fact-checking using 4,500 claims spanning nine languages across six language families. Besides building a systematic performance-evaluation pipeline covering instruction following, authenticity classification, evidence generation, and checking-worthiness scoring, we quantified inequality using standard deviation, coefficient of variation, Gini coefficient, and Theil index. Results showed substantial cross-language disparities, with higher performance on claims from rich-resource languages. To mitigate inequality, we tested two interventions, role-restricted prompt engineering and model fine-tuning. Both approaches reduced disparities, with fine-tuning achieving the largest and most consistent improvement across languages, particularly in checking-worthiness scoring. This study provides a reproducible framework for quantifying multilingual performance and fairness in LLM-based fact-checking and offers practical guidance for developing more equitable verification systems across diverse linguistic contexts.

1 citation

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1016/j.ipm.2026.104616

Or copy a formatted citation

@article{dandan2026,
  title        = {{Performance unfairness of large language models in cross-language fact-checking}},
  author       = {Dandan Wang et al.},
  journal      = {Information Processing and Management},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1016/j.ipm.2026.104616},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Performance unfairness of large language models in cross-language fact-checking

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.37

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.16 × 0.4 = 0.06
M · momentum0.53 × 0.15 = 0.08
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.