Comparison of the Usability and Reliability of Answers to Clinical Questions: AI-Generated ChatGPT versus a Human-Authored Resource

被引：3

作者：

Manian, Farrin A. ^{[1
]}

Garland, Katherine ^{[1
]}

Ding, Jimin ^{[2
]}

机构：

[1] Mercy Hosp St Louis, Dept Med, 3019 Doctors Bldg,Tower B, St Louis, MO 63141 USA

[2] Washington Univ, Dept Stat & Data Sci, St Louis, MO USA

来源：

SOUTHERN MEDICAL JOURNAL | 2024年 / 117卷 / 08期

关键词：

artificial intelligence; ChatGPT; Pearls4Peers; reliability; usability;

D O I：

10.14423/SMJ.0000000000001715

中图分类号：

R5 [内科学];

学科分类号：

1002 ; 100201 ;

摘要：

Objectives: Our aim was to compare the usability and reliability of answers to clinical questions posed of Chat-Generative Pre-Trained Transformer (ChatGPT) compared to those of a human-authored Web source (www.Pearls4Peers.com) in response to "real-world" clinical questions raised during the care of patients. Methods: Two domains of clinical information quality were studied: usability, based on organization/readability, relevance, and usefulness, and reliability, based on clarity, accuracy, and thoroughness. The top 36 most viewed real-world questions from a human-authored Web site (www.Pearls4Peers.com [P4P]) were posed to ChatGPT 3.5. Anonymized answers by ChatGPT and P4P (without literature citations) were separately assessed for usability by 18 practicing physicians ("clinician users") in triplicate and for reliability by 21 expert providers ("content experts") on a Likert scale ("definitely yes," "generally yes," or "no") in duplicate or triplicate. Participants also directly compared the usability and reliability of paired answers. Results: The usability and reliability of ChatGPT answers varied widely depending on the question posed. ChatGPT answers were not considered useful or accurate in 13.9% and 13.1% of cases, respectively. In within-individual rankings for usability, ChatGPT was inferior to P4P in organization/readability, relevance, and usefulness in 29.6%, 28.3%, and 29.6% of cases, respectively, and for reliability, inferior to P4P in clarity, accuracy, and thoroughness in 38.1%, 34.5%, and 31% of cases, respectively. Conclusions: The quality of ChatGPT responses to real-world clinical questions varied widely, with nearly one-third or more answers considered inferior to a human-authored source in several aspects of usability and reliability. Caution is advised when using ChatGPT in clinical decision making.

引用

页码：467 / 473

页数：7

共 16 条

[1] Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum [J].

Ayers, John W. ;

Poliak, Adam ;

Dredze, Mark ;

Leas, Eric C. ;

Zhu, Zechariah ;

Kelley, Jessica B. ;

Faix, Dennis J. ;

Goodman, Aaron M. ;

Longhurst, Christopher A. ;

Hogarth, Michael ;

Smith, Davey M. .

JAMA INTERNAL MEDICINE, 2023, 183 (06) :589-596

[2] CONSTRUCTING CONFIDENCE SETS USING RANK STATISTICS [J].

BAUER, DF .

JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION, 1972, 67 (339) :687-690

[3] NOTE ON AN EXACT TREATMENT OF CONTINGENCY, GOODNESS OF FIT AND OTHER PROBLEMS OF SIGNIFICANCE [J].

FREEMAN, GH ;

HALTON, JH .

BIOMETRIKA, 1951, 38 (1-2) :141-149

[4] How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment [J].

Gilson, Aidan ;

Safranek, Conrad W. ;

Huang, Thomas ;

Socrates, Vimig ;

Chi, Ling ;

Taylor, Richard Andrew ;

Chartash, David .

JMIR MEDICAL EDUCATION, 2023, 9

[5] AI-Generated Medical Advice-GPT and Beyond [J].

Haupt, Claudia E. ;

Marks, Mason .

JAMA-JOURNAL OF THE AMERICAN MEDICAL ASSOCIATION, 2023, 329 (16) :1349-1350

[6]

Hollander M., 1999, NONPARAMETRIC STAT M, V2nd

[7]

Howard G., 2011, Alternation, V4, P288

[8]

Johnson Douglas, 2023, Res Sq, DOI 10.21203/rs.3.rs-2566942/v1

[9] Using ChatGPT to evaluate cancer myths and misconceptions: artificial intelligence and cancer information [J].

Johnson, Skyler B. ;

King, Andy J. ;

Warner, Echo L. ;

Aneja, Sanjay ;

Kann, Benjamin H. ;

Bylund, Carma L. .

JNCI CANCER SPECTRUM, 2023, 7 (02)

[10] Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models [J].

Kung, Tiffany H. ;

Cheatham, Morgan ;

Medenilla, Arielle ;

Sillos, Czarina ;

De Leon, Lorie ;

Elepano, Camille ;

Madriaga, Maria ;

Aggabao, Rimel ;

Diaz-Candido, Giezel ;

Maningo, James ;

Tseng, Victor .

PLOS DIGITAL HEALTH, 2023, 2 (02)

← 1 2 →