• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

生成式人工智能能否提供准确的医疗建议?:以ChatGPT与神经外科医师协会急性颈椎和脊髓损伤临床指南管理为例

Can generative artificial intelligence provide accurate medical advice?: a case of ChatGPT versus Congress of Neurological Surgeons management of acute cervical spine and spinal cord injuries clinical guidelines.

作者信息

Saturno Michael, Mejia Mateo Restrepo, Ahmed Wasil, Yu Alexander, Duey Akiro, Zaidat Bashar, Hijji Fady, Markowitz Jonathan, Kim Jun, Cho Samuel

机构信息

Icahn School of Medicine at Mount Sinai, New York, NY, USA.

出版信息

Asian Spine J. 2025 Mar 4. doi: 10.31616/asj.2024.0301.

DOI:10.31616/asj.2024.0301
PMID:40033723
Abstract

STUDY DESIGN

An experimental study.

PURPOSE

To explore the concordance of ChatGPT responses with established national guidelines for the management of cervical spine and spinal cord injuries.

OVERVIEW OF LITERATURE

ChatGPT-4.0 is an artificial intelligence model that can synthesize large volumes of data and may provide surgeons with recommendations for the management of spinal cord injuries. However, no available literature has quantified ChatGPT's capacity to provide accurate recommendations for the management of cervical spine and spinal cord injuries.

METHODS

Referencing the "Management of acute cervical spine and spinal cord injuries" guidelines published by the Congress of Neurological Surgeons (CNS), a total of 36 questions were formulated. Questions were stratified into therapeutic, diagnostic, or clinical assessment categories as seen in the guidelines. Questions were secondarily grouped according to whether the corresponding recommendation contained level I evidence (highest quality) versus only level II/III evidence (moderate and low quality). ChatGPT-4.0 was prompted with each question, and its responses were assessed by two independent reviewers as "concordant" or "nonconcordant" with the CNS clinical guidelines. "Nonconcordant" responses were rationalized into "insufficient" and "contradictory" categories.

RESULTS

In this study, 22/36 (61.1%) of ChatGPT's responses were concordant with the CNS guidelines. ChatGPT's responses aligned with 17/24 (70.8%) therapeutic questions and 4/7 (57.1%) diagnostic questions. ChatGPT's response aligned with only one of the five clinical assessment questions. Notably, the recommendations supported by level I evidence were the least likely to be replicated by ChatGPT. ChatGPT's responses agreed with 80.8% of the recommendations supported exclusively by level II/III evidence.

CONCLUSIONS

ChatGPT-4 was moderately accurate when generating recommendations that aligned with the clinical guidelines. The model frequently aligned with low evidence and therapeutic recommendations but exhibited inferior performance on topics that contained high-quality evidence or pertained to diagnostic and clinical assessment strategies. Medical practitioners should monitor its usage until further models can be rigorously trained on medical data.

摘要

研究设计

一项实验性研究。

目的

探讨ChatGPT的回答与已确立的颈椎和脊髓损伤管理国家指南的一致性。

文献综述

ChatGPT-4.0是一种人工智能模型,能够综合大量数据,可能为外科医生提供脊髓损伤管理的建议。然而,尚无文献对ChatGPT为颈椎和脊髓损伤管理提供准确建议的能力进行量化。

方法

参照神经外科医师大会(CNS)发布的《急性颈椎和脊髓损伤的管理》指南,共提出36个问题。问题按照指南中的治疗、诊断或临床评估类别进行分层。问题再根据相应建议是否包含I级证据(最高质量)与仅II/III级证据(中等和低质量)进行二次分组。向ChatGPT-4.0提出每个问题,其回答由两名独立评审员评估为与CNS临床指南“一致”或“不一致”。“不一致”的回答被归类为“不充分”和“矛盾”类别。

结果

在本研究中,ChatGPT的22/36(61.1%)回答与CNS指南一致。ChatGPT的回答与17/24(70.8%)的治疗问题和4/7(57.1%)的诊断问题一致。ChatGPT的回答仅与五个临床评估问题中的一个一致。值得注意的是,由I级证据支持的建议最不可能被ChatGPT重复。ChatGPT的回答与仅由II/III级证据支持的80.8%的建议一致。

结论

ChatGPT-4在生成与临床指南一致的建议时具有中等准确性。该模型经常与低证据和治疗建议一致,但在包含高质量证据或涉及诊断和临床评估策略的主题上表现较差。在能够基于医学数据进行严格训练的进一步模型出现之前,医学从业者应监测其使用情况。

相似文献

1
Can generative artificial intelligence provide accurate medical advice?: a case of ChatGPT versus Congress of Neurological Surgeons management of acute cervical spine and spinal cord injuries clinical guidelines.生成式人工智能能否提供准确的医疗建议?:以ChatGPT与神经外科医师协会急性颈椎和脊髓损伤临床指南管理为例
Asian Spine J. 2025 Mar 4. doi: 10.31616/asj.2024.0301.
2
"Dr. AI Will See You Now": How Do ChatGPT-4 Treatment Recommendations Align With Orthopaedic Clinical Practice Guidelines?“AI 医生为您服务”:ChatGPT-4 的治疗建议与骨科临床实践指南如何契合?
Clin Orthop Relat Res. 2024 Dec 1;482(12):2098-2106. doi: 10.1097/CORR.0000000000003234. Epub 2024 Sep 6.
3
Using Artificial Intelligence ChatGPT to Access Medical Information about Chemical Eye Injuries: A Comparative Study.使用人工智能ChatGPT获取有关化学性眼外伤的医学信息:一项比较研究。
JMIR Form Res. 2025 Jun 30. doi: 10.2196/73642.
4
Can generative artificial intelligence pass the orthopaedic board examination?生成式人工智能能通过骨科医师资格考试吗?
J Orthop. 2023 Nov 5;53:27-33. doi: 10.1016/j.jor.2023.10.026. eCollection 2024 Jul.
5
[Volume and health outcomes: evidence from systematic reviews and from evaluation of Italian hospital data].[容量与健康结果:来自系统评价和意大利医院数据评估的证据]
Epidemiol Prev. 2013 Mar-Jun;37(2-3 Suppl 2):1-100.
6
Management of faecal incontinence and constipation in adults with central neurological diseases.成人中枢神经系统疾病患者大便失禁和便秘的管理
Cochrane Database Syst Rev. 2013 Dec 18(12):CD002115. doi: 10.1002/14651858.CD002115.pub4.
7
ChatGPT as a Decision Support Tool in the Management of Chiari I Malformation: A Comparison to 2023 CNS Guidelines.ChatGPT作为Chiari I型畸形管理中的决策支持工具:与2023年中枢神经系统指南的比较
World Neurosurg. 2024 Nov;191:e304-e332. doi: 10.1016/j.wneu.2024.08.122. Epub 2024 Aug 28.
8
Chat Generative Pretraining Transformer Answers Patient-focused Questions in Cervical Spine Surgery.ChatGPT 生成式预训练转换器可回答颈椎手术患者关注的问题。
Clin Spine Surg. 2024 Jul 1;37(6):E278-E281. doi: 10.1097/BSD.0000000000001600. Epub 2024 Mar 21.
9
Artificial Intelligence in Orthopaedics: Performance of ChatGPT on Text and Image Questions on a Complete AAOS Orthopaedic In-Training Examination (OITE).人工智能在骨科领域的应用:ChatGPT 在 AAOS 骨科住院医师培训考试(OITE)全题文本和图像问题上的表现。
J Surg Educ. 2024 Nov;81(11):1645-1649. doi: 10.1016/j.jsurg.2024.08.002. Epub 2024 Sep 14.
10
Artificial Intelligence in Peripheral Artery Disease Education: A Battle Between ChatGPT and Google Gemini.外周动脉疾病教育中的人工智能:ChatGPT与谷歌Gemini的较量
Cureus. 2025 Jun 1;17(6):e85174. doi: 10.7759/cureus.85174. eCollection 2025 Jun.

本文引用的文献

1
Can Large Language Models (LLMs) Predict the Appropriate Treatment of Acute Hip Fractures in Older Adults? Comparing Appropriate Use Criteria With Recommendations From ChatGPT.大语言模型(LLMs)能否预测老年人急性髋部骨折的适当治疗方法?比较适当使用标准与 ChatGPT 的建议
J Am Acad Orthop Surg Glob Res Rev. 2024 Aug 9;8(8). doi: 10.5435/JAAOSGlobal-D-24-00206. eCollection 2024 Aug 1.
2
An analysis of ChatGPT recommendations for the diagnosis and treatment of cervical radiculopathy.对 ChatGPT 推荐的颈神经根病诊断和治疗方案的分析。
J Neurosurg Spine. 2024 Jun 28;41(3):385-395. doi: 10.3171/2024.4.SPINE231148. Print 2024 Sep 1.
3
Performance of a Large Language Model in the Generation of Clinical Guidelines for Antibiotic Prophylaxis in Spine Surgery.
大型语言模型在生成脊柱手术抗生素预防临床指南方面的表现。
Neurospine. 2024 Mar;21(1):128-146. doi: 10.14245/ns.2347310.655. Epub 2024 Mar 31.
4
ChatGPT versus NASS clinical guidelines for degenerative spondylolisthesis: a comparative analysis.ChatGPT 与 NASS 退行性脊柱滑脱临床指南比较分析。
Eur Spine J. 2024 Nov;33(11):4182-4203. doi: 10.1007/s00586-024-08198-6. Epub 2024 Mar 15.
5
Adherence of a Large Language Model to Clinical Guidelines for Craniofacial Plastic and Reconstructive Surgeries.大型语言模型对颅面整形与重建手术临床指南的遵循情况。
Ann Plast Surg. 2024 Mar 1;92(3):261-262. doi: 10.1097/SAP.0000000000003757. Epub 2024 Jan 6.
6
Use of ChatGPT for Determining Clinical and Surgical Treatment of Lumbar Disc Herniation With Radiculopathy: A North American Spine Society Guideline Comparison.使用ChatGPT确定伴神经根病的腰椎间盘突出症的临床和手术治疗:与北美脊柱协会指南的比较
Neurospine. 2024 Mar;21(1):149-158. doi: 10.14245/ns.2347052.526. Epub 2024 Jan 31.
7
Performance of ChatGPT on NASS Clinical Guidelines for the Diagnosis and Treatment of Low Back Pain: A Comparison Study.ChatGPT 在 NASS 腰痛诊断和治疗临床指南中的表现:一项对比研究。
Spine (Phila Pa 1976). 2024 May 1;49(9):640-651. doi: 10.1097/BRS.0000000000004915. Epub 2024 Jan 12.
8
Generative artificial intelligence fails to provide sufficiently accurate recommendations when compared to established breast reconstruction surgery guidelines.与既定的乳房重建手术指南相比,生成式人工智能未能提供足够准确的建议。
J Plast Reconstr Aesthet Surg. 2023 Nov;86:248-250. doi: 10.1016/j.bjps.2023.09.030. Epub 2023 Sep 15.
9
Comparing the Efficacy of Long Spinal Board, Sked Stretcher, and Vacuum Mattress in Cervical Spine Immobilization; a Method-Oriented Experimental Study.比较长脊柱板、Sked担架和真空床垫在颈椎固定中的效果:一项以方法为导向的实验研究。
Arch Acad Emerg Med. 2023 Jun 12;11(1):e44. doi: 10.22037/aaem.v11i1.2036. eCollection 2023.
10
Comparison of Ophthalmologist and Large Language Model Chatbot Responses to Online Patient Eye Care Questions.眼科医生与大型语言模型聊天机器人对在线患者眼部护理问题的回复比较。
JAMA Netw Open. 2023 Aug 1;6(8):e2330320. doi: 10.1001/jamanetworkopen.2023.30320.