ChatGPT Across Domains: A Systematic Review of Applications, Evaluation Approaches, and Open Challenges

Shirin Abbasi & Amir Masoud Rahmani

Expert Systems2026https://doi.org/10.1111/exsy.70244article
AJG 2
Weight
0.50

What the paper says

In recent years, there has been a rise in the use of ChatGPT for education, healthcare, smart cities and emerging technologies. However, no available studies or reviews have provided a consistent and reliable depiction of the situation regarding its usage and evaluation. The reporting of datasets, evaluation indicators, factors influencing performance and conditions of deployment has also varied from study to study. This fragmented state of affairs seriously inhibits attempts to assess ChatGPT's capabilities and limitations and thus improve the design of future versions. Earlier reviews were often conducted in a way pertaining to a single area or were mainly descriptive, with less emphasis on methodological evaluation and issues in deployment and ethics. To fill this void, we undertook a systematic review, according to PRISMA guidelines, limiting our searches to English‐language journal articles published during 2021–2025 by reputable publishers. Such studies focused directly on GPT models, provided assessment conditions and were cited extensively as preprints, while the exclusion criteria encompassed poorly linked studies, those not in English and studies employing ChatGPT as an adjunct. An analysis of these studies revealed that the vast majority of research has taken place in the area of education (32%) and health (28%). The review revealed significant variation in assessment accuracy across domains, frequent challenges with doubtful sensitivity, unpredictable and rapid changes and risks associated with specific domains that impact reliability and safety. This study's primary contribution is an effort to develop an integrated analytical framework that puts together these interdisciplinary results in a streamlined manner for interpreting the capabilities and limitations of ChatGPT. Because of the methodological heterogeneity of existing studies, the results can be viewed as qualitative trends instead of standard quantitative evidence. The results thereby accentuate the need for consistent criteria, domain‐informed evaluation practices and stronger methodological reporting to underpin a more reliable deployment of ChatGPT‐based systems.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1111/exsy.70244

Or copy a formatted citation

@article{shirin2026,
  title        = {{ChatGPT Across Domains: A Systematic Review of Applications, Evaluation Approaches, and Open Challenges}},
  author       = {Shirin Abbasi & Amir Masoud Rahmani},
  journal      = {Expert Systems},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1111/exsy.70244},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

ChatGPT Across Domains: A Systematic Review of Applications, Evaluation Approaches, and Open Challenges

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.