Diagnostic accuracy of large language models in psychiatry

被引：1

作者：

Gargari, Omid Kohandel ^{[1
]}

Fatehi, Farhad ^{[2
,3
]}

Mohammadi, Ida ^{[1
]}

Firouzabadi, Shahryar Rajai ^{[1
]}

Shafiee, Arman ^{[1
]}

Habibi, Gholamreza ^{[1
]}

机构：

[1] Farzan Clin Res Inst, Farzan Artificial Intelligence Team, Tehran, Iran

[2] Univ Queensland, Fac Med, Ctr Hlth Serv Res, Brisbane, Australia

[3] Monash Univ, Sch Psychol Sci, Melbourne, Australia

来源：

ASIAN JOURNAL OF PSYCHIATRY | 2024年 / 100卷

关键词：

Artificial intelligence (AI); Psychiatry; Diagnostic accuracy; Large Language Models (LLMs); DSM-5 clinical vignettes; Natural Language Processing (NLP); CARDIOVASCULAR-DISEASES; ARTIFICIAL-INTELLIGENCE; DECISION-MAKING; PREDICTION;

D O I：

10.1016/j.ajp.2024.104168

中图分类号：

R749 [精神病学];

学科分类号：

100205 ;

摘要：

Introduction: Medical decision-making is crucial for effective treatment, especially in psychiatry where diagnosis often relies on subjective patient reports and a lack of high-specificity symptoms. Artificial intelligence (AI), particularly Large Language Models (LLMs) like GPT, has emerged as a promising tool to enhance diagnostic accuracy in psychiatry. This comparative study explores the diagnostic capabilities of several AI models, including Aya, GPT-3.5, GPT-4, GPT-3.5 clinical assistant (CA), Nemotron, and Nemotron CA, using clinical cases from the DSM-5. Methods: We curated 20 clinical cases from the DSM-5 Clinical Cases book, covering a wide range of psychiatric diagnoses. Four advanced AI models (GPT-3.5 Turbo, GPT-4, Aya, Nemotron) were tested using prompts to elicit detailed diagnoses and reasoning. The models' performances were evaluated based on accuracy and quality of reasoning, with additional analysis using the Retrieval Augmented Generation (RAG) methodology for models accessing the DSM-5 text. Results: The AI models showed varied diagnostic accuracy, with GPT-3.5 and GPT-4 performing notably better than Aya and Nemotron in terms of both accuracy and reasoning quality. While models struggled with specific disorders such as cyclothymic and disruptive mood dysregulation disorders, others excelled, particularly in diagnosing psychotic and bipolar disorders. Statistical analysis highlighted significant differences in accuracy and reasoning, emphasizing the superiority of the GPT models. Discussion: The application of AI in psychiatry offers potential improvements in diagnostic accuracy. The superior performance of the GPT models can be attributed to their advanced natural language processing capabilities and extensive training on diverse text data, enabling more effective interpretation of psychiatric language. However, models like Aya and Nemotron showed limitations in reasoning, indicating a need for further refinement in their training and application. Conclusion: AI holds significant promise for enhancing psychiatric diagnostics, with certain models demonstrating high potential in interpreting complex clinical descriptions accurately. Future research should focus on expanding the dataset and integrating multimodal data to further enhance the diagnostic capabilities of AI in psychiatry.

引用

页数：6

共 50 条

[41] A Surgical Perspective on Large Language Models
Miller, Robert
ANNALS OF SURGERY, 2023, 278 (02) : E211 - E213
[42] Large Language Models in Cosmetic Dermatology
Landau, Marina
Kroumpouzos, George
Goldust, Mohamad
JOURNAL OF COSMETIC DERMATOLOGY, 2025, 24 (02)
[43] Large Language Models for Cultural Heritage
Trichopoulos, Georgios
2ND INTERNATIONAL CONFERENCE OF THE GREECE ACM SIGCHI CHAPTER, CHIGREECE 2023, 2023,
[44] Large language models for medicine: a survey
Zheng, Yanxin
Gan, Wensheng
Chen, Zefeng
Qi, Zhenlian
Liang, Qian
Yu, Philip S.
INTERNATIONAL JOURNAL OF MACHINE LEARNING AND CYBERNETICS, 2025, 16 (02) : 1015 - 1040
[45] Meaning and understanding in large language models
Havlik, Vladimir
SYNTHESE, 2024, 205 (01)
[46] Applications of Large Language Models in Pathology
Cheng, Jerome
BIOENGINEERING-BASEL, 2024, 11 (04):
[47] Applications of large language models in oncology
Loefflert, Chiara M.
Bressem, Keno K.
Truhn, Daniel
ONKOLOGIE, 2024, 30 (05): : 388 - 393
[48] Adversarial Attacks on Large Language Models
Zou, Jing
Zhang, Shungeng
Qiu, Meikang
KNOWLEDGE SCIENCE, ENGINEERING AND MANAGEMENT, PT IV, KSEM 2024, 2024, 14887 : 85 - 96
[49] Quo Vadis ChatGPT? From large language models to Large Knowledge Models
Venkatasubramanian, Venkat
Chakraborty, Arijit
COMPUTERS & CHEMICAL ENGINEERING, 2025, 192
[50] The Language of Creativity: Evidence from Humans and Large Language Models
Orwig, William
Edenbaum, Emma R.
Greene, Joshua D.
Schacter, Daniel L.
JOURNAL OF CREATIVE BEHAVIOR, 2024, 58 (01): : 128 - 136

← 1 2 3 4 5 →