Inferential performance and temporal stability of large language models in suicide method prediction: A forensic psychiatric analysis

Halit Canberk Aydoğan et al.

Health Informatics Journal2026https://doi.org/10.1177/14604582251414578article
ABDC C
Weight
0.50

What the paper says

<b>Objective:</b> This study presents a structured evaluation of large language models (LLMs) in predicting suicide methods based exclusively on indirect forensic psychiatric indicators. <b>Methods:</b> Ninety-two forensic psychiatric cases (2019-2024), involving survivors of suicide attempts formally examined in medico-legal contexts, were retrospectively analyzed. Variables included age, sex, psychiatric diagnosis, previous suicide attempts, psychiatric medication use, impulsivity, and consciousness at emergency admission. Six LLMs were tested: ChatGPT-4o, ChatGPT-4o Mini, ChatGPT-O3 (OpenAI), Gemini 2.0 Flash, Gemini 2.5 Pro, and Gemini 2.5 Flash (Google DeepMind). Each case was converted into a standardized anonymized prompt. Model predictions were categorized by blinded forensic physicians and evaluated using accuracy, precision, recall, F1-score, and Cohen's Kappa for 1-month reproducibility. <b>Results:</b> Gemini 2.5 Flash achieved the highest performance with 76.09% accuracy, 46.9% F1-score, and 45.2% recall. It accurately predicted the dominant method, medication overdose, but underperformed for rare categories. Temporal reproducibility was moderate (κ = 0.582), while other models exhibited lower and less stable performance. <b>Conclusion:</b> LLMs can infer suicide methods from indirect psychiatric data with encouraging accuracy. However, limitations in detecting rare methods and maintaining temporal consistency suggest the need for further methodological refinement and external validation prior to forensic application.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1177/14604582251414578

Or copy a formatted citation

@article{halit2026,
  title        = {{Inferential performance and temporal stability of large language models in suicide method prediction: A forensic psychiatric analysis}},
  author       = {Halit Canberk Aydoğan et al.},
  journal      = {Health Informatics Journal},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1177/14604582251414578},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Inferential performance and temporal stability of large language models in suicide method prediction: A forensic psychiatric analysis

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.