通过对临床文档的自然语言处理识别痴呆患者的功能状态障碍：横断面研究。

Identifying Functional Status Impairment in People Living With Dementia Through Natural Language Processing of Clinical Documents: Cross-Sectional Study.

机构信息

Department of Medicine, Brigham and Women's Hospital, Boston, MA, United States.

Harvard Medical School, Boston, MA, United States.

出版信息

J Med Internet Res. 2024 Feb 13;26:e47739. doi: 10.2196/47739.

DOI:10.2196/47739

PMID:38349732

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC10900085/

Abstract

BACKGROUND

Assessment of activities of daily living (ADLs) and instrumental ADLs (iADLs) is key to determining the severity of dementia and care needs among older adults. However, such information is often only documented in free-text clinical notes within the electronic health record and can be challenging to find.

OBJECTIVE

This study aims to develop and validate machine learning models to determine the status of ADL and iADL impairments based on clinical notes.

METHODS

This cross-sectional study leveraged electronic health record clinical notes from Mass General Brigham's Research Patient Data Repository linked with Medicare fee-for-service claims data from 2007 to 2017 to identify individuals aged 65 years or older with at least 1 diagnosis of dementia. Notes for encounters both 180 days before and after the first date of dementia diagnosis were randomly sampled. Models were trained and validated using note sentences filtered by expert-curated keywords (filtered cohort) and further evaluated using unfiltered sentences (unfiltered cohort). The model's performance was compared using area under the receiver operating characteristic curve and area under the precision-recall curve (AUPRC).

RESULTS

The study included 10,000 key-term-filtered sentences representing 441 people (n=283, 64.2% women; mean age 82.7, SD 7.9 years) and 1000 unfiltered sentences representing 80 people (n=56, 70% women; mean age 82.8, SD 7.5 years). Area under the receiver operating characteristic curve was high for the best-performing ADL and iADL models on both cohorts (>0.97). For ADL impairment identification, the random forest model achieved the best AUPRC (0.89, 95% CI 0.86-0.91) on the filtered cohort; the support vector machine model achieved the highest AUPRC (0.82, 95% CI 0.75-0.89) for the unfiltered cohort. For iADL impairment, the Bio+Clinical bidirectional encoder representations from transformers (BERT) model had the highest AUPRC (filtered: 0.76, 95% CI 0.68-0.82; unfiltered: 0.58, 95% CI 0.001-1.0). Compared with a keyword-search approach on the unfiltered cohort, machine learning reduced false-positive rates from 4.5% to 0.2% for ADL and 1.8% to 0.1% for iADL.

CONCLUSIONS

In this study, we demonstrated the ability of machine learning models to accurately identify ADL and iADL impairment based on free-text clinical notes, which could be useful in determining the severity of dementia.

摘要

背景

评估日常生活活动（ADL）和工具性日常生活活动（iADL）是确定老年人痴呆严重程度和护理需求的关键。然而，这些信息通常仅记录在电子健康记录中的自由文本临床记录中，并且难以找到。

目的

本研究旨在开发和验证机器学习模型，以基于临床记录确定 ADL 和 iADL 损伤的状态。

方法

这项横断面研究利用了马萨诸塞州综合医院 Brigham 的研究患者数据存储库中的电子健康记录临床记录，并与 2007 年至 2017 年的医疗保险按服务收费数据相关联，以确定至少有 1 次痴呆诊断的 65 岁及以上的个体。从痴呆诊断的第一个日期之前和之后的 180 天随机抽取就诊记录。使用经过专家精心策划的关键字（过滤队列）过滤的注释句子来训练和验证模型，并使用未过滤的句子（未过滤队列）进一步评估模型。使用接收器操作特征曲线下面积和精度-召回曲线下面积（AUPRC）比较模型的性能。

结果

该研究包括 10000 个关键词过滤句子，代表 441 人（n=283，64.2%为女性；平均年龄 82.7，标准差 7.9 岁）和 1000 个未过滤句子，代表 80 人（n=56，70%为女性；平均年龄 82.8，标准差 7.5 岁）。最佳 ADL 和 iADL 模型在两个队列上的接收器操作特征曲线下面积均较高（>0.97）。对于 ADL 损伤识别，随机森林模型在过滤队列中实现了最佳的 AUPRC（0.89，95%CI 0.86-0.91）；支持向量机模型在未过滤队列中实现了最高的 AUPRC（0.82，95%CI 0.75-0.89）。对于 iADL 损伤，基于生物+临床双向转换器表示（BERT）的模型具有最高的 AUPRC（过滤：0.76，95%CI 0.68-0.82；未过滤：0.58，95%CI 0.001-1.0）。与未过滤队列上的关键字搜索方法相比，机器学习将 ADL 的假阳性率从 4.5%降低到 0.2%，将 iADL 的假阳性率从 1.8%降低到 0.1%。

结论

在这项研究中，我们证明了机器学习模型基于自由文本临床记录准确识别 ADL 和 iADL 损伤的能力，这可能有助于确定痴呆的严重程度。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/6b4f/10900085/f891ea9f5551/jmir_v26i1e47739_fig1.jpg

相似文献

Identifying Functional Status Impairment in People Living With Dementia Through Natural Language Processing of Clinical Documents: Cross-Sectional Study.通过对临床文档的自然语言处理识别痴呆患者的功能状态障碍：横断面研究。

J Med Internet Res. 2024 Feb 13;26:e47739. doi: 10.2196/47739.

Signs and symptoms to determine if a patient presenting in primary care or hospital outpatient settings has COVID-19.在基层医疗机构或医院门诊环境中，如果患者出现以下症状和体征，可判断其是否患有 COVID-19。

Cochrane Database Syst Rev. 2022 May 20;5(5):CD013665. doi: 10.1002/14651858.CD013665.pub3.

Comparison of Two Modern Survival Prediction Tools, SORG-MLA and METSSS, in Patients With Symptomatic Long-bone Metastases Who Underwent Local Treatment With Surgery Followed by Radiotherapy and With Radiotherapy Alone.两种现代生存预测工具 SORG-MLA 和 METSSS 在接受手术联合放疗和单纯放疗治疗有症状长骨转移患者中的比较。

Clin Orthop Relat Res. 2024 Dec 1;482(12):2193-2208. doi: 10.1097/CORR.0000000000003185. Epub 2024 Jul 23.

Are Current Survival Prediction Tools Useful When Treating Subsequent Skeletal-related Events From Bone Metastases?当前的生存预测工具在治疗骨转移后的骨骼相关事件时有用吗？

Clin Orthop Relat Res. 2024 Sep 1;482(9):1710-1721. doi: 10.1097/CORR.0000000000003030. Epub 2024 Mar 22.

A New Measure of Quantified Social Health Is Associated With Levels of Discomfort, Capability, and Mental and General Health Among Patients Seeking Musculoskeletal Specialty Care.一种新的量化社会健康指标与寻求肌肉骨骼专科护理的患者的不适程度、能力以及心理和总体健康水平相关。

Clin Orthop Relat Res. 2025 Apr 1;483(4):647-663. doi: 10.1097/CORR.0000000000003394. Epub 2025 Feb 5.

Clinical judgement by primary care physicians for the diagnosis of all-cause dementia or cognitive impairment in symptomatic people.初级保健医生对有症状人群进行全因痴呆或认知障碍诊断的临床判断。

Cochrane Database Syst Rev. 2022 Jun 16;6(6):CD012558. doi: 10.1002/14651858.CD012558.pub2.

Natural Language Processing of Clinical Documentation to Assess Functional Status in Patients With Heart Failure.临床文档的自然语言处理用于评估心力衰竭患者的功能状态。

JAMA Netw Open. 2024 Nov 4;7(11):e2443925. doi: 10.1001/jamanetworkopen.2024.43925.

Occupational therapy for cognitive impairment in stroke patients.脑卒中患者认知障碍的作业治疗。

Cochrane Database Syst Rev. 2022 Mar 29;3(3):CD006430. doi: 10.1002/14651858.CD006430.pub3.

Falls prevention interventions for community-dwelling older adults: systematic review and meta-analysis of benefits, harms, and patient values and preferences.社区居住的老年人跌倒预防干预措施：系统评价和荟萃分析的益处、危害以及患者的价值观和偏好。

Syst Rev. 2024 Nov 26;13(1):289. doi: 10.1186/s13643-024-02681-3.

Extraction of sleep information from clinical notes of Alzheimer's disease patients using natural language processing.使用自然语言处理从阿尔茨海默病患者的临床记录中提取睡眠信息。

J Am Med Inform Assoc. 2024 Oct 1;31(10):2217-2227. doi: 10.1093/jamia/ocae177.

引用本文的文献

Dual-stream algorithms for dementia detection: Harnessing structured and unstructured electronic health record data, a novel approach to prevalence estimation.用于痴呆症检测的双流算法：利用结构化和非结构化电子健康记录数据，一种估计患病率的新方法。

Alzheimers Dement. 2025 May;21(5):e70132. doi: 10.1002/alz.70132.

Identifying Patient-Reported Outcome Measure Documentation in Veterans Health Administration Chiropractic Clinic Notes: Natural Language Processing Analysis.识别退伍军人健康管理局脊椎按摩诊所记录中的患者报告结局测量文档：自然语言处理分析

JMIR Med Inform. 2025 Apr 2;13:e66466. doi: 10.2196/66466.

Natural language processing in Alzheimer's disease research: Systematic review of methods, data, and efficacy.阿尔茨海默病研究中的自然语言处理：方法、数据和疗效的系统综述

Alzheimers Dement (Amst). 2025 Feb 11;17(1):e70082. doi: 10.1002/dad2.70082. eCollection 2025 Jan-Mar.

本文引用的文献

2023 Alzheimer's disease facts and figures.2023 年阿尔茨海默病事实和数据。

Alzheimers Dement. 2023 Apr;19(4):1598-1695. doi: 10.1002/alz.13016. Epub 2023 Mar 14.

Development and Validation of a Deep Learning Model for Detection of Allergic Reactions Using Safety Event Reports Across Hospitals.利用医院安全事件报告开发和验证一种用于检测过敏反应的深度学习模型

JAMA Netw Open. 2020 Nov 2;3(11):e2022836. doi: 10.1001/jamanetworkopen.2020.22836.

Natural Language Processing Reveals Vulnerable Mental Health Support Groups and Heightened Health Anxiety on Reddit During COVID-19: Observational Study.自然语言处理揭示了新冠疫情期间Reddit上脆弱的心理健康支持小组以及加剧的健康焦虑：一项观察性研究。

J Med Internet Res. 2020 Oct 12;22(10):e22635. doi: 10.2196/22635.

Why Do Older Korean Adults Respond Differently to Activities of Daily Living and Instrumental Activities of Daily Living? A Differential Item Functioning Analysis.为什么韩国老年成年人对日常生活活动和工具性日常生活活动的反应不同？一项差异项目功能分析。

Ann Geriatr Med Res. 2019 Dec;23(4):197-203. doi: 10.4235/agmr.19.0047. Epub 2019 Dec 30.

Impact of Instrumental Activities of Daily Living Limitations on Hospital Readmission: an Observational Study Using Machine Learning.日常生活工具性活动受限对医院再入院的影响：一项使用机器学习的观察性研究

J Gen Intern Med. 2020 Oct;35(10):2865-2872. doi: 10.1007/s11606-020-05982-0. Epub 2020 Jul 29.

Function-based dementia severity assessment for vascular cognitive impairment.基于功能的血管性认知障碍严重程度评估。

J Formos Med Assoc. 2021 Jan;120(1 Pt 2):533-541. doi: 10.1016/j.jfma.2020.07.001. Epub 2020 Jul 9.

External Validation of an Algorithm to Identify Patients with High Data-Completeness in Electronic Health Records for Comparative Effectiveness Research.一种用于比较效果研究的识别电子健康记录中数据完整性高的患者的算法的外部验证

Clin Epidemiol. 2020 Feb 4;12:133-141. doi: 10.2147/CLEP.S232540. eCollection 2020.

BioBERT: a pre-trained biomedical language representation model for biomedical text mining.BioBERT：一种用于生物医学文本挖掘的预训练生物医学语言表示模型。

Bioinformatics. 2020 Feb 15;36(4):1234-1240. doi: 10.1093/bioinformatics/btz682.

Use of Natural Language Processing to Extract Clinical Cancer Phenotypes from Electronic Medical Records.利用自然语言处理从电子病历中提取临床癌症表型

Cancer Res. 2019 Nov 1;79(21):5463-5470. doi: 10.1158/0008-5472.CAN-19-0579. Epub 2019 Aug 8.

The Value of Unstructured Electronic Health Record Data in Geriatric Syndrome Case Identification.非结构化电子健康记录数据在老年综合征病例识别中的价值。

J Am Geriatr Soc. 2018 Aug;66(8):1499-1507. doi: 10.1111/jgs.15411. Epub 2018 Jul 4.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

通过对临床文档的自然语言处理识别痴呆患者的功能状态障碍：横断面研究。

Identifying Functional Status Impairment in People Living With Dementia Through Natural Language Processing of Clinical Documents: Cross-Sectional Study.

机构信息

出版信息

BACKGROUND

OBJECTIVE

METHODS

RESULTS

CONCLUSIONS

背景

目的

方法

结果

结论

相似文献

引用本文的文献

本文引用的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献