聊天机器人在牙髓病临床决策支持中的比较性能：一项为期4天的准确性和一致性研究。

Comparative Performance of Chatbots in Endodontic Clinical Decision Support: A 4-Day Accuracy and Consistency Study.

作者信息

Büker Mine, Sümbüllü Meltem, Arslan Hakan

机构信息

Faculty of Dentistry, Department of Endodontics, Mersin University, Mersin, Turkey.

Faculty of Dentistry, Department of Endodontics, Atatürk University, Erzurum, Turkey.

出版信息

Int Dent J. 2025 Jul 27;75(5):100920. doi: 10.1016/j.identj.2025.100920.

DOI:10.1016/j.identj.2025.100920

PMID:40720933

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC12319550/

Abstract

INTRODUCTION AND AIMS

Despite the use of artificial intelligence, which is increasingly prevalent in healthcare settings, concerns remain regarding its reliability and accuracy. The study assessed the overall, difficulty level-specific, and day-to-day accuracy and consistency of 5 AI chatbots-ChatGPT-3.5, ChatGPT-4.o, Gemini 2.0 Flash, Copilot, and Copilot Pro-in answering clinically relevant endodontic questions.

METHODS

Seventy-six correct/incorrect questions were developed by 2 endodontists and categorized by an expert into 3 difficulty levels (Basic [B]-, Intermediate [I]-, and Advanced [A]- level]. Twenty questions from each difficulty level were selected from a set of 74 validated questions (B, n = 26; I, n = 24; A, n = 24), resulting in a total of 60 questions. The questions were asked of the chatbots over a period of 4 days, at 3 different times each day (morning, afternoon, and evening).

RESULTS

ChatGPT-4.o achieved the highest overall accuracy (82.5%) and perfect performance in the B-level category (95.0%), while Copilot Pro had the lowest accuracy (74.03%). Gemini and ChatGPT-3.5 showed similar overall accuracy. Gemini's accuracy significantly improved over time, whereas significant decreases were noted in the Copilot Pro model across days, and no significant change was detected in both ChatGPT models and Copilot. In the B-level category, while Copilot Pro showed a significant decrease in accuracy rates, and in the B- and I-level categories, Copilot showed a significant increase in accuracy rates over the days. In the A-level category, Gemini demonstrated a significant increase in accuracy rates over the days.

CONCLUSIONS

ChatGPT-4.o demonstrated superior performance, whereas Copilot and Copilot Pro showed insufficient accuracy. ChatGPT-3.5 and Gemini may be acceptable for general queries but require caution in more advanced cases.

CLINICAL RELEVANCE

ChatGPT-4.o demonstrated the highest overall accuracy and consistency in all question categories over 4 days, suggesting its potential as a reliable tool for clinical decision-making.

摘要

引言与目的

尽管人工智能在医疗环境中的应用日益普遍，但人们仍对其可靠性和准确性存在担忧。本研究评估了5个人工智能聊天机器人——ChatGPT-3.5、ChatGPT-4.0、Gemini 2.0 Flash、Copilot和Copilot Pro——在回答临床相关牙髓病问题时的总体、特定难度级别以及日常的准确性和一致性。

方法

2名牙髓病医生编写了76道正误问题，并由一名专家将其分为3个难度级别（基础[B]级、中级[I]级和高级[A]级）。从一组74道经过验证的问题（B级，n = 26；I级，n = 24；A级，n = 24）中，每个难度级别选取20个问题，共60个问题。在4天的时间里，每天3个不同时间（上午、下午和晚上）向聊天机器人提问这些问题。

结果

ChatGPT-4.0的总体准确率最高（82.5%），在B级类别中表现完美（95.0%），而Copilot Pro的准确率最低（74.03%）。Gemini和ChatGPT-3.5的总体准确率相似。Gemini的准确率随时间显著提高，而Copilot Pro模型在不同日期的准确率显著下降，ChatGPT两个模型和Copilot均未检测到显著变化。在B级类别中，Copilot Pro的准确率显著下降，在B级和I级类别中，Copilot的准确率在不同日期显著提高。在A级类别中，Gemini在不同日期的准确率显著提高。

结论

ChatGPT-4.0表现出卓越的性能，而Copilot和Copilot Pro的准确率不足。ChatGPT-3.5和Gemini对于一般查询可能是可接受的，但在更复杂的情况下需要谨慎使用。

临床相关性

ChatGPT-4.0在4天内所有问题类别中表现出最高的总体准确率和一致性，表明其作为临床决策可靠工具的潜力。

相似文献

Comparative Performance of Chatbots in Endodontic Clinical Decision Support: A 4-Day Accuracy and Consistency Study.聊天机器人在牙髓病临床决策支持中的比较性能：一项为期4天的准确性和一致性研究。

Int Dent J. 2025 Jul 27;75(5):100920. doi: 10.1016/j.identj.2025.100920.

Accuracy and Reliability of Artificial Intelligence Chatbots as Public Information Sources in Implant Dentistry.人工智能聊天机器人作为种植牙科公共信息来源的准确性和可靠性

Int J Oral Maxillofac Implants. 2025 Jun 25;0(0):1-23. doi: 10.11607/jomi.11280.

Accuracy of ChatGPT-3.5, ChatGPT-4o, Copilot, Gemini, Claude, and Perplexity in advising on lumbosacral radicular pain against clinical practice guidelines: cross-sectional study.ChatGPT-3.5、ChatGPT-4o、Copilot、Gemini、Claude和Perplexity在依据临床实践指南对腰骶神经根性疼痛提供建议方面的准确性：横断面研究

Front Digit Health. 2025 Jun 27;7:1574287. doi: 10.3389/fdgth.2025.1574287. eCollection 2025.

Accuracy of ChatGPT, Gemini, Copilot, and Claude to Blepharoplasty-Related Questions.ChatGPT、Gemini、Copilot和Claude对双眼皮手术相关问题的回答准确性。

Aesthetic Plast Surg. 2025 Jul 21. doi: 10.1007/s00266-025-05071-9.

Performance of 3 Conversational Generative Artificial Intelligence Models for Computing Maximum Safe Doses of Local Anesthetics: Comparative Analysis.用于计算局部麻醉药最大安全剂量的3种对话式生成人工智能模型的性能：比较分析

JMIR AI. 2025 May 13;4:e66796. doi: 10.2196/66796.

Artificial Intelligence in Peripheral Artery Disease Education: A Battle Between ChatGPT and Google Gemini.外周动脉疾病教育中的人工智能：ChatGPT与谷歌Gemini的较量

Cureus. 2025 Jun 1;17(6):e85174. doi: 10.7759/cureus.85174. eCollection 2025 Jun.

Benchmarking AI Chatbots for Maternal Lactation Support: A Cross-Platform Evaluation of Quality, Readability, and Clinical Accuracy.用于产妇泌乳支持的人工智能聊天机器人基准测试：质量、可读性和临床准确性的跨平台评估

Healthcare (Basel). 2025 Jul 20;13(14):1756. doi: 10.3390/healthcare13141756.

Evaluating large language models for renal colic imaging recommendations: a comparative analysis of Gemini, copilot, and ChatGPT-4.0.评估用于肾绞痛成像建议的大语言模型：Gemini、Copilot和ChatGPT-4.0的比较分析。

Int J Emerg Med. 2025 Jul 4;18(1):123. doi: 10.1186/s12245-025-00895-3.

Comparison of artificial intelligence systems in answering prosthodontics questions from the dental specialty exam in Turkey.土耳其牙科专业考试中人工智能系统回答口腔修复学问题的比较

J Dent Sci. 2025 Jul;20(3):1454-1459. doi: 10.1016/j.jds.2025.01.025. Epub 2025 Jan 31.

Performance of 7 Artificial Intelligence Chatbots on Board-style Endodontic Questions.7款人工智能聊天机器人在根管治疗式问题上的表现

J Endod. 2025 Jun 26. doi: 10.1016/j.joen.2025.06.014.

本文引用的文献

The Transformative Role of Artificial Intelligence in Dentistry: A Comprehensive Overview. Part 1: Fundamentals of AI, and its Contemporary Applications in Dentistry.人工智能在牙科领域的变革性作用：全面概述。第1部分：人工智能基础及其在牙科领域的当代应用。

Int Dent J. 2025 Apr;75(2):383-396. doi: 10.1016/j.identj.2025.02.005. Epub 2025 Mar 11.

Artificial intelligence chatbots in transfusion medicine: A cross-sectional study.输血医学中的人工智能聊天机器人：一项横断面研究。

Vox Sang. 2025 Mar 5. doi: 10.1111/vox.70009.

The Transformative Role of Artificial Intelligence in Dentistry: A Comprehensive Overview Part 2: The Promise and Perils, and the International Dental Federation Communique.人工智能在牙科领域的变革性作用：全面概述第2部分：前景与风险，以及国际牙科联合会公报

Int Dent J. 2025 Apr;75(2):397-404. doi: 10.1016/j.identj.2025.02.006. Epub 2025 Feb 25.

Evaluation of Chatbots in the Emergency Management of Avulsion Injuries.聊天机器人在撕脱伤应急管理中的评估

Dent Traumatol. 2025 Aug;41(4):437-444. doi: 10.1111/edt.13041. Epub 2025 Jan 24.

Comparative accuracy of artificial intelligence chatbots in pulpal and periradicular diagnosis: A cross-sectional study.人工智能聊天机器人在牙髓和根尖周诊断中的比较准确性：一项横断面研究。

Comput Biol Med. 2024 Dec;183:109332. doi: 10.1016/j.compbiomed.2024.109332. Epub 2024 Oct 30.

Assessment of artificial intelligence applications in responding to dental trauma.评估人工智能在应对牙科创伤中的应用。

Dent Traumatol. 2024 Dec;40(6):722-729. doi: 10.1111/edt.12965. Epub 2024 May 14.

Artificial Intelligence (AI)-driven dental education: Exploring the role of chatbots in a clinical learning environment.人工智能驱动的牙科教育：探索聊天机器人在临床学习环境中的作用。

J Prosthet Dent. 2024 Apr 20. doi: 10.1016/j.prosdent.2024.03.038.

Evidence-based potential of generative artificial intelligence large language models in orthodontics: a comparative study of ChatGPT, Google Bard, and Microsoft Bing.生成式人工智能大语言模型在正畸学中的循证潜力：ChatGPT、谷歌巴德和微软必应的比较研究

Eur J Orthod. 2024 Apr 13. doi: 10.1093/ejo/cjae017.

Artificial intelligence chatbots and large language models in dental education: Worldwide survey of educators.人工智能聊天机器人和大型语言模型在口腔医学教育中的应用：教育者的全球调查。

Eur J Dent Educ. 2024 Nov;28(4):865-876. doi: 10.1111/eje.13009. Epub 2024 Apr 8.

Accuracy and consistency of chatbots versus clinicians for answering pediatric dentistry questions: A pilot study.聊天机器人与临床医生回答儿科牙科问题的准确性和一致性：一项试点研究。

J Dent. 2024 May;144:104938. doi: 10.1016/j.jdent.2024.104938. Epub 2024 Apr 3.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

聊天机器人在牙髓病临床决策支持中的比较性能：一项为期4天的准确性和一致性研究。

Comparative Performance of Chatbots in Endodontic Clinical Decision Support: A 4-Day Accuracy and Consistency Study.

作者信息

机构信息

出版信息

INTRODUCTION AND AIMS

METHODS

RESULTS

CONCLUSIONS

CLINICAL RELEVANCE

引言与目的

方法

结果

结论

临床相关性

相似文献

本文引用的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

本文引用的文献