挑战聊天机器人：对 ChatGPT 对 DBP 病例研究的诊断和建议的评估。

Challenging the Chatbot: An Assessment of ChatGPT's Diagnoses and Recommendations for DBP Case Studies.

机构信息

Division of Developmental and Behavioral Pediatrics, Steven and Alexandra Cohen's Children Medical Center of New York, Lake Success, NY.

出版信息

J Dev Behav Pediatr. 2024 Jan 1;45(1):e8-e13. doi: 10.1097/DBP.0000000000001255. Epub 2024 Feb 9.

DOI:10.1097/DBP.0000000000001255

PMID:38347665

Abstract

OBJECTIVE

Chat Generative Pretrained Transformer-3.5 (ChatGPT) is a publicly available and free artificial intelligence chatbot that logs billions of visits per day; parents may rely on such tools for developmental and behavioral medical consultations. The objective of this study was to determine how ChatGPT evaluates developmental and behavioral pediatrics (DBP) case studies and makes recommendations and diagnoses.

METHODS

ChatGPT was asked to list treatment recommendations and a diagnosis for each of 97 DBP case studies. A panel of 3 DBP physicians evaluated ChatGPT's diagnostic accuracy and scored treatment recommendations on accuracy (5-point Likert scale) and completeness (3-point Likert scale). Physicians also assessed whether ChatGPT's treatment plan correctly addressed cultural and ethical issues for relevant cases. Scores were analyzed using Python, and descriptive statistics were computed.

RESULTS

The DBP panel agreed with ChatGPT's diagnosis for 66.2% of the case reports. The mean accuracy score of ChatGPT's treatment plan was deemed by physicians to be 4.6 (between entirely correct and more correct than incorrect), and the mean completeness was 2.6 (between complete and adequate). Physicians agreed that ChatGPT addressed relevant cultural issues in 10 out of the 11 appropriate cases and the ethical issues in the single ethical case.

CONCLUSION

While ChatGPT can generate a comprehensive and adequate list of recommendations, the diagnosis accuracy rate is still low. Physicians must advise caution to patients when using such online sources.

摘要

目的

Chat Generative Pretrained Transformer-3.5（ChatGPT）是一个公开的、免费的人工智能聊天机器人，每天有数以亿计的访问量；家长们可能会依赖这些工具进行发育和行为医学咨询。本研究的目的是确定 ChatGPT 如何评估发育和行为儿科学（DBP）病例，并提出建议和诊断。

方法

我们要求 ChatGPT 为 97 个 DBP 病例列出治疗建议和诊断。一个 DBP 医生小组评估了 ChatGPT 的诊断准确性，并对治疗建议的准确性（5 分李克特量表）和完整性（3 分李克特量表）进行评分。医生们还评估了 ChatGPT 的治疗计划是否正确地解决了相关病例的文化和伦理问题。使用 Python 分析评分，并计算描述性统计数据。

结果

DBP 小组同意 ChatGPT 对 66.2%的病例报告的诊断。医生认为 ChatGPT 的治疗方案的平均准确性得分为 4.6（介于完全正确和更正确之间），平均完整性得分为 2.6（介于完整和足够之间）。医生们一致认为，ChatGPT 在 11 个适当的案例中解决了 10 个相关的文化问题，在单一的伦理案例中解决了伦理问题。