Does language bias GenAI academic evaluation in humanities and social sciences? A mixed‐methods study based on Chinese‐language HSS papers

Yu Zhu et al.

Journal of the Association for Information Science and Technology (JASIST)2026https://doi.org/10.1002/asi.70079article
AJG 3
Weight
0.50

What the paper says

As generative AI (GenAI) systems are increasingly deployed in cross‐language research evaluation, whether GenAI evaluates multilingual scholarship without language‐induced bias remains unclear. This study examines language bias patterns in GenAI evaluation of humanities and social sciences (HSS) research across models and disciplines. Using a within‐subjects design, 1150 expert‐selected papers from 23 disciplines were evaluated by GPT‐4o and DeepSeek‐V3 in Chinese and English. Results reveal opposite language biases depending on model type: GPT‐4o favors English (Cohen's d = 1.10), while DeepSeek‐V3 favors Chinese (Cohen's d = −0.87), persisting across all disciplines. Thematic analysis reveals a systematic decoupling between scores and evaluative reasoning: both models generate more critical comments for English papers, yet arrive at opposite scores through different rhetorical strategies—GPT‐4o tends to moderate its positive assessments of Chinese papers while DeepSeek‐V3 amplifies them. This decoupling suggests that bias is embedded in the multi‐layered pathways through which models generate and aggregate evaluations. This study provides controlled evidence that language bias in GenAI evaluation is bidirectional and model‐dependent, with scores not directly reflecting evaluative justifications. The findings have implications for designing fairer multilingual academic evaluation systems and for understanding the limitations of GenAI as scholarly evaluation infrastructure.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1002/asi.70079

Or copy a formatted citation

@article{yu2026,
  title        = {{Does language bias GenAI academic evaluation in humanities and social sciences? A mixed‐methods study based on Chinese‐language HSS papers}},
  author       = {Yu Zhu et al.},
  journal      = {Journal of the Association for Information Science and Technology (JASIST)},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1002/asi.70079},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Does language bias GenAI academic evaluation in humanities and social sciences? A mixed‐methods study based on Chinese‐language HSS papers

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.