Can Machines Think Like Humans: A Behavioral Evaluation of LLM Agents in Dictator Games

Ji Ma

Voluntas: International Journal of Voluntary and Non-Profit Organizations2026https://doi.org/10.1017/s0957876526000173article
AJG 2ABDC B
Weight
0.50

What the paper says

Abstract As large language model (LLM)-based agents increasingly engage with human society, how well do we understand their prosocial behaviors? We (1) investigate how LLM agents’ prosocial behaviors can be induced by different personas and benchmarked against human behaviors and (2) introduce a social science approach to evaluate LLM agents’ decision-making. We explored how different personas and experimental framings affect these AI agents’ altruistic behavior in dictator games and compared their behaviors within the same LLM family, across various families, and with human behaviors. The findings reveal that merely assigning a human-like identity to LLMs does not produce human-like behaviors. They suggest that LLM agents’ reasoning does not consistently exhibit textual markers of human decision-making in dictator games and that their alignment with human behavior varies substantially across model architectures and prompt formulations; even worse, such dependence does not follow a clear pattern. As society increasingly integrates machine intelligence, “prosocial AI” emerges as a promising and urgent research direction in philanthropic studies.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1017/s0957876526000173

Or copy a formatted citation

@article{ji2026,
  title        = {{Can Machines Think Like Humans: A Behavioral Evaluation of LLM Agents in Dictator Games}},
  author       = {Ji Ma},
  journal      = {Voluntas: International Journal of Voluntary and Non-Profit Organizations},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1017/s0957876526000173},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Can Machines Think Like Humans: A Behavioral Evaluation of LLM Agents in Dictator Games

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.