Evaluating AI in medicine: a comparative analysis of expert and ChatGPT responses to colorectal cancer questions

被引:16
作者
Peng, Wen [1 ,2 ]
Feng, Yifei [1 ,2 ]
Yao, Cui [1 ,2 ]
Zhang, Sheng [3 ]
Zhuo, Han [4 ]
Qiu, Tianzhu [5 ]
Zhang, Yi [1 ,2 ]
Tang, Junwei [1 ,2 ]
Gu, Yanhong [5 ]
Sun, Yueming [1 ,2 ]
机构
[1] Nanjing Med Univ, Affiliated Hosp 1, Dept Gen Surg, Nanjing 210029, Jiangsu, Peoples R China
[2] Nanjing Med Univ, Sch Clin Med 1, Nanjing, Peoples R China
[3] Nanjing Med Univ, Affiliated Hosp 1, Dept Radiotherapy, Nanjing, Peoples R China
[4] Nanjing Med Univ, Affiliated Hosp 1, Dept Intervent, Nanjing, Peoples R China
[5] Nanjing Med Univ, Affiliated Hosp 1, Dept Oncol, Nanjing, Peoples R China
关键词
D O I
10.1038/s41598-024-52853-3
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
Colorectal cancer (CRC) is a global health challenge, and patient education plays a crucial role in its early detection and treatment. Despite progress in AI technology, as exemplified by transformer-like models such as ChatGPT, there remains a lack of in-depth understanding of their efficacy for medical purposes. We aimed to assess the proficiency of ChatGPT in the field of popular science, specifically in answering questions related to CRC diagnosis and treatment, using the book "Colorectal Cancer: Your Questions Answered" as a reference. In general, 131 valid questions from the book were manually input into ChatGPT. Responses were evaluated by clinical physicians in the relevant fields based on comprehensiveness and accuracy of information, and scores were standardized for comparison. Not surprisingly, ChatGPT showed high reproducibility in its responses, with high uniformity in comprehensiveness, accuracy, and final scores. However, the mean scores of ChatGPT's responses were significantly lower than the benchmarks, indicating it has not reached an expert level of competence in CRC. While it could provide accurate information, it lacked in comprehensiveness. Notably, ChatGPT performed well in domains of radiation therapy, interventional therapy, stoma care, venous care, and pain control, almost rivaling the benchmarks, but fell short in basic information, surgery, and internal medicine domains. While ChatGPT demonstrated promise in specific domains, its general efficiency in providing CRC information falls short of expert standards, indicating the need for further advancements and improvements in AI technology for patient education in healthcare.
引用
收藏
页数:16
相关论文
共 20 条
  • [1] [Anonymous], 2023, ChatGPT-Release Notes
  • [2] Misinformation on the Internet regarding Ablative Therapies for Prostate Cancer
    Asafu-Adjei, Denise
    Mikkilineni, Nina
    Sebesta, Elisabeth
    Hyams, Elias
    [J]. UROLOGY, 2019, 133 : 182 - 186
  • [3] Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum
    Ayers, John W.
    Poliak, Adam
    Dredze, Mark
    Leas, Eric C.
    Zhu, Zechariah
    Kelley, Jessica B.
    Faix, Dennis J.
    Goodman, Aaron M.
    Longhurst, Christopher A.
    Hogarth, Michael
    Smith, Davey M.
    [J]. JAMA INTERNAL MEDICINE, 2023, 183 (06) : 589 - 596
  • [4] Therapeutic landscape and future direction of metastatic colorectal cancer
    Bando, Hideaki
    Ohtsu, Atsushi
    Yoshino, Takayuki
    [J]. NATURE REVIEWS GASTROENTEROLOGY & HEPATOLOGY, 2023, 20 (05) : 306 - 322
  • [5] Role of Chat GPT in Public Health
    Biswas, Som S.
    [J]. ANNALS OF BIOMEDICAL ENGINEERING, 2023, 51 (05) : 868 - 869
  • [6] Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios
    Cascella, Marco
    Montomoli, Jonathan
    Bellini, Valentina
    Bignami, Elena
    [J]. JOURNAL OF MEDICAL SYSTEMS, 2023, 47 (01)
  • [7] Gu Y., 2019, Colorectal Cancer: Your Questions Answered
  • [8] High-quality health systems in the Sustainable Development Goals era: time for a revolution
    Kruk, Margaret E.
    Gage, Anna D.
    Arsenault, Catherine
    Jordan, Keely
    Leslie, Hannah H.
    Roder-DeWan, Sanam
    Adeyi, Olusoji
    Barker, Pierre
    Daelmans, Bernadette
    Doubova, Svetlana V.
    English, Mike
    Garcia Elorrio, Ezequiel
    Guanais, Frederico
    Gureje, Oye
    Hirschhorn, Lisa R.
    Jiang, Lixin
    Kelley, Edward
    Lemango, Ephrem Tekle
    Liljestrand, Jerker
    Malata, Address
    Marchant, Tanya
    Matsoso, Malebona Precious
    Meara, John G.
    Mohanan, Manoj
    Ndiaye, Youssoupha
    Norheim, Ole F.
    Reddy, K. Srinath
    Rowe, Alexander K.
    Salomon, Joshua A.
    Thapa, Gagan
    Twum-Danso, Nana A. Y.
    Pate, Muhammad
    [J]. LANCET GLOBAL HEALTH, 2018, 6 (11): : E1196 - E1252
  • [9] Colorectal cancer burden, trends and risk factors in China: A review and comparison with the United States
    Li, Qianru
    Wu, Hongliang
    Cao, Maomao
    Li, He
    He, Siyi
    Yang, Fan
    Yan, Xinxin
    Zhang, Shaoli
    Teng, Yi
    Xia, Changfa
    Peng, Ji
    Chen, Wanqing
    [J]. CHINESE JOURNAL OF CANCER RESEARCH, 2022, 34 (05) : 483 - +
  • [10] Cancer prevention and screening: the next step in the era of precision medicine
    Loomans-Kropp, Holli A.
    Umar, Asad
    [J]. NPJ PRECISION ONCOLOGY, 2019, 3 (1)