• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

编辑评论:现成的大语言模型质量不足以提供医疗建议,而定制大语言模型则能产生高质量的建议。

Editorial Commentary: Off-the-Shelf Large Language Models Are of Insufficient Quality to Provide Medical Treatment Recommendations, While Customization of Large Language Models Results in Quality Recommendations.

作者信息

Ramkumar Prem N, Masotto Andrew F, Woo Joshua J

机构信息

Commons Clinic (A.F.M., J.J.W.); The Warren Alpert Medical School of Brown University (J.J.W.).

出版信息

Arthroscopy. 2025 Feb;41(2):276-278. doi: 10.1016/j.arthro.2024.09.047. Epub 2024 Oct 3.

DOI:10.1016/j.arthro.2024.09.047
PMID:39368620
Abstract

The content accuracy of off-the-shelf large language models (LLMs) mirrors the content accuracy of the unregulated Internet from which these generative artificial intelligence models are supplied. With error rates approximating 30% in terms of treatment recommendations for the management of common musculoskeletal conditions, seeking expert opinion remains paramount. However, custom LLMs represent an excellent opportunity to infuse niche, bespoke expertise from the many specialties and subspecialties within medicine. Methods of customizing these generative models broadly fall under the categories of prompt engineering; "retrieval-augmented generation" prioritizing retrieval of relevant information from a specific domain of data; "fine-tuning" of a basic pretrained model into one that is refined for health care-related vernacular and acronyms; and "agentic augmentation" including software that breaks down complex tasks into smaller ones, recruiting multiple LLMs (with or without retrieval-augmented generation), optimizing the output, internally deciding whether the response is appropriate or sufficient, and even passing on an unmet outcome to a human for supervision ("phone a friend"). Custom LLMs offer physicians and their associated organizations the rare opportunity to regain control of our profession by re-establishing authority in our increasingly digital landscape.

摘要

现成的大语言模型(LLMs)的内容准确性反映了这些生成式人工智能模型所基于的无监管互联网的内容准确性。在常见肌肉骨骼疾病管理的治疗建议方面,错误率接近30%,因此寻求专家意见仍然至关重要。然而,定制大语言模型是一个绝佳机会,可以融入医学众多专业和亚专业领域的细分、定制化专业知识。定制这些生成式模型的方法大致可分为以下几类:提示工程;“检索增强生成”,即优先从特定数据领域检索相关信息;将基本的预训练模型“微调”为针对医疗保健相关术语和首字母缩略词进行优化的模型;以及“智能体增强”,包括将复杂任务分解为较小任务的软件,调用多个大语言模型(有无检索增强生成均可),优化输出,内部判断响应是否合适或充分,甚至将未解决的结果传递给人类进行监督(“求助热线”)。定制大语言模型为医生及其相关组织提供了一个难得的机会,通过在日益数字化的环境中重新确立权威来重新掌控我们的职业。

相似文献

1
Editorial Commentary: Off-the-Shelf Large Language Models Are of Insufficient Quality to Provide Medical Treatment Recommendations, While Customization of Large Language Models Results in Quality Recommendations.编辑评论:现成的大语言模型质量不足以提供医疗建议,而定制大语言模型则能产生高质量的建议。
Arthroscopy. 2025 Feb;41(2):276-278. doi: 10.1016/j.arthro.2024.09.047. Epub 2024 Oct 3.
2
Large Language Models Applied to Health Care Tasks May Improve Clinical Efficiency, Value of Care Rendered, Research, and Medical Education.应用于医疗保健任务的大语言模型可能会提高临床效率、所提供护理的价值、研究水平以及医学教育质量。
Arthroscopy. 2025 Mar;41(3):547-556. doi: 10.1016/j.arthro.2024.12.010. Epub 2024 Dec 16.
3
Custom Large Language Models Improve Accuracy: Comparing Retrieval Augmented Generation and Artificial Intelligence Agents to Noncustom Models for Evidence-Based Medicine.定制大语言模型提高准确性:将检索增强生成和人工智能代理与非定制模型在循证医学方面进行比较
Arthroscopy. 2025 Mar;41(3):565-573.e6. doi: 10.1016/j.arthro.2024.10.042. Epub 2024 Nov 7.
4
Unveiling the Potential of Large Language Models in Transforming Chronic Disease Management: Mixed Methods Systematic Review.揭示大语言模型在转变慢性病管理中的潜力:混合方法系统评价
J Med Internet Res. 2025 Apr 16;27:e70535. doi: 10.2196/70535.
5
Large Language Models for Therapy Recommendations Across 3 Clinical Specialties: Comparative Study.大型语言模型在 3 个临床专业领域的治疗推荐中的应用:比较研究。
J Med Internet Res. 2023 Oct 30;25:e49324. doi: 10.2196/49324.
6
Currently Available Large Language Models Do Not Provide Musculoskeletal Treatment Recommendations That Are Concordant With Evidence-Based Clinical Practice Guidelines.目前可用的大语言模型并未提供与循证临床实践指南相一致的肌肉骨骼治疗建议。
Arthroscopy. 2025 Feb;41(2):263-275.e6. doi: 10.1016/j.arthro.2024.07.040. Epub 2024 Aug 22.
7
Quality of Answers of Generative Large Language Models Versus Peer Users for Interpreting Laboratory Test Results for Lay Patients: Evaluation Study.生成式大语言模型与同行用户对解释非专业患者实验室检测结果的答案质量比较:评估研究。
J Med Internet Res. 2024 Apr 17;26:e56655. doi: 10.2196/56655.
8
Examining the Role of Large Language Models in Orthopedics: Systematic Review.检查大型语言模型在骨科中的作用:系统评价。
J Med Internet Res. 2024 Nov 15;26:e59607. doi: 10.2196/59607.
9
Utilizing large language models for gastroenterology research: a conceptual framework.利用大语言模型进行胃肠病学研究:一个概念框架。
Therap Adv Gastroenterol. 2025 Apr 1;18:17562848251328577. doi: 10.1177/17562848251328577. eCollection 2025.
10
Evaluating and Enhancing Japanese Large Language Models for Genetic Counseling Support: Comparative Study of Domain Adaptation and the Development of an Expert-Evaluated Dataset.评估和增强用于遗传咨询支持的日本大语言模型:领域适应的比较研究与专家评估数据集的开发
JMIR Med Inform. 2025 Jan 16;13:e65047. doi: 10.2196/65047.