• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

评估大型语言模型用于起草急诊科就诊总结。

Evaluating large language models for drafting emergency department encounter summaries.

作者信息

Williams Christopher Y K, Bains Jaskaran, Tang Tianyu, Patel Kishan, Lucas Alexa N, Chen Fiona, Miao Brenda Y, Butte Atul J, Kornblith Aaron E

机构信息

Bakar Computational Health Sciences Institute, University of California, San Francisco, California, United States of America.

Department of Emergency Medicine, University of California, San Francisco, California, United States of America.

出版信息

PLOS Digit Health. 2025 Jun 17;4(6):e0000899. doi: 10.1371/journal.pdig.0000899. eCollection 2025 Jun.

DOI:10.1371/journal.pdig.0000899
PMID:40526634
原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC12173386/
Abstract

Large language models (LLMs) possess a range of capabilities which may be applied to the clinical domain, including text summarization. As ambient artificial intelligence scribes and other LLM-based tools begin to be deployed within healthcare settings, rigorous evaluations of the accuracy of these technologies are urgently needed. In this cross-sectional study of 100 randomly sampled adult Emergency Department (ED) visits from 2012 to 2023 at the University of California, San Francisco ED, we sought to investigate the performance of GPT-4 and GPT-3.5-turbo in generating ED encounter summaries and evaluate the prevalence and type of errors for each section of the encounter summary across three evaluation criteria: 1) Inaccuracy of LLM-summarized information; 2) Hallucination of information; 3) Omission of relevant clinical information. In total, 33% of summaries generated by GPT-4 and 10% of those generated by GPT-3.5-turbo were entirely error-free across all evaluated domains. Summaries generated by GPT-4 were mostly accurate, with inaccuracies found in only 10% of cases, however, 42% of the summaries exhibited hallucinations and 47% omitted clinically relevant information. Inaccuracies and hallucinations were most commonly found in the Plan sections of LLM-generated summaries, while clinical omissions were concentrated in text describing patients' Physical Examination findings or History of Presenting Complaint. The potential harmfulness score across errors was low, with a mean score of 0.57 (SD 1.11) out of 7 and only three errors scoring 4 ('Potential for permanent harm') or greater. In summary, we found that LLMs could generate accurate encounter summaries but were liable to hallucination and omission of clinically relevant information. Individual errors on average had a low potential for harm. A comprehensive understanding of the location and type of errors found in LLM-generated clinical text is important to facilitate clinician review of such content and prevent patient harm.

摘要

大语言模型(LLMs)具备一系列可应用于临床领域的能力,包括文本摘要。随着环境人工智能抄写员和其他基于大语言模型的工具开始在医疗环境中部署,迫切需要对这些技术的准确性进行严格评估。在这项横断面研究中,我们从2012年至2023年在加利福尼亚大学旧金山分校急诊科随机抽取了100例成年患者的急诊就诊病例,旨在研究GPT - 4和GPT - 3.5 - turbo生成急诊就诊摘要的性能,并根据三个评估标准评估就诊摘要各部分错误的发生率和类型:1)大语言模型总结信息的不准确;2)信息幻觉;3)相关临床信息的遗漏。总体而言,GPT - 4生成的摘要中有33%在所有评估领域完全无错误,GPT - 3.5 - turbo生成的摘要中有10%完全无错误。GPT - 4生成的摘要大多准确,只有10%的病例存在不准确情况,然而,42%的摘要出现了幻觉,47%遗漏了临床相关信息。不准确和幻觉最常见于大语言模型生成的摘要的“计划”部分,而临床信息遗漏集中在描述患者体格检查结果或现病史的文本中。错误的潜在危害评分较低,7分制的平均得分为0.57(标准差1.11),只有三个错误的评分达到了4分(“有永久伤害的可能性”)或更高。总之,我们发现大语言模型可以生成准确的就诊摘要,但容易出现幻觉和遗漏临床相关信息。平均而言,个别错误造成伤害的可能性较低。全面了解大语言模型生成的临床文本中错误的位置和类型,对于促进临床医生对此类内容的审查以及防止患者受到伤害非常重要。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a500/12173386/21495e7afde7/pdig.0000899.g003.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a500/12173386/f3aeb464ecc3/pdig.0000899.g001.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a500/12173386/c51cab323946/pdig.0000899.g002.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a500/12173386/21495e7afde7/pdig.0000899.g003.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a500/12173386/f3aeb464ecc3/pdig.0000899.g001.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a500/12173386/c51cab323946/pdig.0000899.g002.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a500/12173386/21495e7afde7/pdig.0000899.g003.jpg

相似文献

1
Evaluating large language models for drafting emergency department encounter summaries.评估大型语言模型用于起草急诊科就诊总结。
PLOS Digit Health. 2025 Jun 17;4(6):e0000899. doi: 10.1371/journal.pdig.0000899. eCollection 2025 Jun.
2
Evaluating Large Language Models for Drafting Emergency Department Discharge Summaries.评估用于起草急诊科出院小结的大语言模型。
medRxiv. 2024 Apr 4:2024.04.03.24305088. doi: 10.1101/2024.04.03.24305088.
3
Signs and symptoms to determine if a patient presenting in primary care or hospital outpatient settings has COVID-19.在基层医疗机构或医院门诊环境中,如果患者出现以下症状和体征,可判断其是否患有 COVID-19。
Cochrane Database Syst Rev. 2022 May 20;5(5):CD013665. doi: 10.1002/14651858.CD013665.pub3.
4
Drugs for preventing postoperative nausea and vomiting in adults after general anaesthesia: a network meta-analysis.成人全身麻醉后预防术后恶心呕吐的药物:网状Meta分析
Cochrane Database Syst Rev. 2020 Oct 19;10(10):CD012859. doi: 10.1002/14651858.CD012859.pub2.
5
Home treatment for mental health problems: a systematic review.心理健康问题的居家治疗:一项系统综述
Health Technol Assess. 2001;5(15):1-139. doi: 10.3310/hta5150.
6
Clinical judgement by primary care physicians for the diagnosis of all-cause dementia or cognitive impairment in symptomatic people.初级保健医生对有症状人群进行全因痴呆或认知障碍诊断的临床判断。
Cochrane Database Syst Rev. 2022 Jun 16;6(6):CD012558. doi: 10.1002/14651858.CD012558.pub2.
7
Behavioral interventions to reduce risk for sexual transmission of HIV among men who have sex with men.降低男男性行为者中艾滋病毒性传播风险的行为干预措施。
Cochrane Database Syst Rev. 2008 Jul 16(3):CD001230. doi: 10.1002/14651858.CD001230.pub2.
8
Sertindole for schizophrenia.用于治疗精神分裂症的舍吲哚。
Cochrane Database Syst Rev. 2005 Jul 20;2005(3):CD001715. doi: 10.1002/14651858.CD001715.pub2.
9
A rapid and systematic review of the clinical effectiveness and cost-effectiveness of paclitaxel, docetaxel, gemcitabine and vinorelbine in non-small-cell lung cancer.对紫杉醇、多西他赛、吉西他滨和长春瑞滨在非小细胞肺癌中的临床疗效和成本效益进行的快速系统评价。
Health Technol Assess. 2001;5(32):1-195. doi: 10.3310/hta5320.
10
Systemic treatments for metastatic cutaneous melanoma.转移性皮肤黑色素瘤的全身治疗
Cochrane Database Syst Rev. 2018 Feb 6;2(2):CD011123. doi: 10.1002/14651858.CD011123.pub2.

引用本文的文献

1
Research progress and implications of the application of large language model in shared decision-making in China's healthcare field.大语言模型在中国医疗领域共享决策应用中的研究进展与启示
Front Public Health. 2025 Jul 10;13:1605212. doi: 10.3389/fpubh.2025.1605212. eCollection 2025.
2
Verifiable Summarization of Electronic Health Records Using Large Language Models to Support Chart Review.使用大语言模型对电子健康记录进行可验证的摘要以支持病历审查。
medRxiv. 2025 Jun 3:2025.06.02.25328807. doi: 10.1101/2025.06.02.25328807.
3
Artificial intelligence-enabled precision medicine for inflammatory skin diseases.

本文引用的文献

1
Physician- and Large Language Model-Generated Hospital Discharge Summaries.医生和大语言模型生成的医院出院小结
JAMA Intern Med. 2025 May 5. doi: 10.1001/jamainternmed.2025.0821.
2
Use of a Large Language Model to Assess Clinical Acuity of Adults in the Emergency Department.使用大型语言模型评估急诊科成人的临床敏锐度。
JAMA Netw Open. 2024 May 1;7(5):e248895. doi: 10.1001/jamanetworkopen.2024.8895.
3
The Limits of Clinician Vigilance as an AI Safety Bulwark.临床医生警觉性作为人工智能安全保障的局限性。
用于炎症性皮肤病的人工智能精准医学。
ArXiv. 2025 May 14:arXiv:2505.09527v1.
JAMA. 2024 Apr 9;331(14):1173-1174. doi: 10.1001/jama.2024.3620.
4
Adapted large language models can outperform medical experts in clinical text summarization.经过改编的大型语言模型在临床文本总结方面的表现优于医学专家。
Nat Med. 2024 Apr;30(4):1134-1142. doi: 10.1038/s41591-024-02855-5. Epub 2024 Feb 27.
5
Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study.评估 GPT-4 在医疗保健中延续种族和性别偏见的潜力:一项模型评估研究。
Lancet Digit Health. 2024 Jan;6(1):e12-e22. doi: 10.1016/S2589-7500(23)00225-X.
6
Will Generative Artificial Intelligence Deliver on Its Promise in Health Care?生成式人工智能能否在医疗保健领域兑现其承诺?
JAMA. 2024 Jan 2;331(1):65-69. doi: 10.1001/jama.2023.25054.
7
Primary Care Physicians' Perspectives on High-Quality Discharge Summaries.基层医疗医生对高质量出院小结的看法。
J Gen Intern Med. 2024 Jun;39(8):1438-1443. doi: 10.1007/s11606-023-08541-5. Epub 2023 Nov 27.
8
Patterns in Physician Burnout in a Stable-Linked Cohort.稳定关联队列中医生倦怠的模式。
JAMA Netw Open. 2023 Oct 2;6(10):e2336745. doi: 10.1001/jamanetworkopen.2023.36745.
9
Potential of ChatGPT and GPT-4 for Data Mining of Free-Text CT Reports on Lung Cancer.ChatGPT 和 GPT-4 在挖掘肺癌 CT 报告自由文本数据方面的潜力
Radiology. 2023 Sep;308(3):e231362. doi: 10.1148/radiol.231362.
10
A method to automate the discharge summary hospital course for neurology patients.一种自动化神经内科患者出院小结住院流程的方法。
J Am Med Inform Assoc. 2023 Nov 17;30(12):1995-2003. doi: 10.1093/jamia/ocad177.