Division of Otolaryngology, Department of Surgical Sciences, Università degli Studi di Torino, Turin, Italy.
Otolaryngology Unit, Santi Paolo e Carlo Hospital, Department of Health Sciences, Università degli Studi di Milano, Milan, Italy.
Eur Arch Otorhinolaryngol. 2024 Sep;281(9):5001-5006. doi: 10.1007/s00405-024-08746-2. Epub 2024 May 25.
This study evaluates the efficacy of two advanced Large Language Models (LLMs), OpenAI's ChatGPT 4 and Google's Gemini Advanced, in providing treatment recommendations for head and neck oncology cases. The aim is to assess their utility in supporting multidisciplinary oncological evaluations and decision-making processes.
This comparative analysis examined the responses of ChatGPT 4 and Gemini Advanced to five hypothetical cases of head and neck cancer, each representing a different anatomical subsite. The responses were evaluated against the latest National Comprehensive Cancer Network (NCCN) guidelines by two blinded panels using the total disagreement score (TDS) and the artificial intelligence performance instrument (AIPI). Statistical assessments were performed using the Wilcoxon signed-rank test and the Friedman test.
Both LLMs produced relevant treatment recommendations with ChatGPT 4 generally outperforming Gemini Advanced regarding adherence to guidelines and comprehensive treatment planning. ChatGPT 4 showed higher AIPI scores (median 3 [2-4]) compared to Gemini Advanced (median 2 [2-3]), indicating better overall performance. Notably, inconsistencies were observed in the management of induction chemotherapy and surgical decisions, such as neck dissection.
While both LLMs demonstrated the potential to aid in the multidisciplinary management of head and neck oncology, discrepancies in certain critical areas highlight the need for further refinement. The study supports the growing role of AI in enhancing clinical decision-making but also emphasizes the necessity for continuous updates and validation against current clinical standards to integrate AI into healthcare practices fully.
本研究评估了两种先进的大型语言模型(LLM),OpenAI 的 ChatGPT 4 和 Google 的 Gemini Advanced,在为头颈部肿瘤病例提供治疗建议方面的效果。旨在评估它们在支持多学科肿瘤评估和决策过程中的效用。
这项对比分析评估了 ChatGPT 4 和 Gemini Advanced 对五个头颈部癌症假设病例的反应,每个病例代表不同的解剖亚部位。通过两个盲法小组使用总分歧评分(TDS)和人工智能绩效工具(AIPI),将这些反应与最新的国家综合癌症网络(NCCN)指南进行评估。使用 Wilcoxon 符号秩检验和 Friedman 检验进行统计评估。
两种 LLM 都提出了相关的治疗建议,ChatGPT 4 在遵守指南和全面治疗计划方面普遍优于 Gemini Advanced。ChatGPT 4 的 AIPI 评分(中位数 3 [2-4])高于 Gemini Advanced(中位数 2 [2-3]),表明总体性能更好。值得注意的是,在诱导化疗和手术决策(如颈部清扫术)的管理方面观察到了不一致性。
虽然两种 LLM 都显示出在头颈部肿瘤多学科管理中辅助的潜力,但在某些关键领域的差异突出表明需要进一步改进。该研究支持人工智能在增强临床决策方面的作用不断增长,但也强调了需要不断更新并针对当前临床标准进行验证,以充分将人工智能整合到医疗保健实践中。