一种使用综合电子健康记录数据进行风险预测的半参数方法。

A SEMIPARAMETRIC METHOD FOR RISK PREDICTION USING INTEGRATED ELECTRONIC HEALTH RECORD DATA.

作者信息

Hasler Jill, Ma Yanyuan, Wei Yizheng, Parikh Ravi, Chen Jinbo

机构信息

Fox Chase Cancer Center.

Department of Statistics, Pennsylvania State University.

出版信息

Ann Appl Stat. 2024 Dec;18(4):3318-3337. doi: 10.1214/24-AOAS1938. Epub 2024 Oct 31.

DOI:10.1214/24-AOAS1938

PMID:40134753

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC11934126/

Abstract

When using electronic health records (EHRs) for clinical and translational research, additional data is often available from external sources to enrich the information extracted from EHRs. For example, academic biobanks have more granular data available, and patient reported data is often collected through small-scale surveys. It is common that the external data is available only for a small subset of patients who have EHR information. We propose efficient and robust methods for building and evaluating models for predicting the risk of binary outcomes using such integrated EHR data. Our method is built upon an idea derived from the two-phase design literature that modeling the availability of a patient's external data as a function of an EHR-based preliminary predictive score leads to effective utilization of the EHR data. Through both theoretical and simulation studies, we show that our method has high efficiency for estimating log-odds ratio parameters, the area under the ROC curve, as well as other measures for quantifying predictive accuracy. We apply our method to develop a model for predicting the short-term mortality risk of oncology patients, where the data was extracted from the University of Pennsylvania hospital system EHR and combined with survey-based patient reported outcome data.

摘要

在将电子健康记录（EHR）用于临床和转化研究时，通常可以从外部来源获取额外数据，以丰富从EHR中提取的信息。例如，学术生物样本库拥有更详细的数据，且患者报告的数据通常通过小规模调查收集。常见的情况是，外部数据仅适用于拥有EHR信息的一小部分患者。我们提出了高效且稳健的方法，用于构建和评估使用此类整合EHR数据预测二元结局风险的模型。我们的方法基于从两阶段设计文献中衍生出的一个想法，即将患者外部数据的可用性建模为基于EHR的初步预测分数的函数，从而实现EHR数据的有效利用。通过理论和模拟研究，我们表明我们的方法在估计对数优势比参数、ROC曲线下面积以及其他量化预测准确性的指标方面具有很高的效率。我们应用我们的方法开发了一个预测肿瘤患者短期死亡风险的模型，该模型的数据从宾夕法尼亚大学医院系统的EHR中提取，并与基于调查的患者报告结局数据相结合。

相似文献

A SEMIPARAMETRIC METHOD FOR RISK PREDICTION USING INTEGRATED ELECTRONIC HEALTH RECORD DATA.一种使用综合电子健康记录数据进行风险预测的半参数方法。

Ann Appl Stat. 2024 Dec;18(4):3318-3337. doi: 10.1214/24-AOAS1938. Epub 2024 Oct 31.

Comparison of Two Modern Survival Prediction Tools, SORG-MLA and METSSS, in Patients With Symptomatic Long-bone Metastases Who Underwent Local Treatment With Surgery Followed by Radiotherapy and With Radiotherapy Alone.两种现代生存预测工具 SORG-MLA 和 METSSS 在接受手术联合放疗和单纯放疗治疗有症状长骨转移患者中的比较。

Clin Orthop Relat Res. 2024 Dec 1;482(12):2193-2208. doi: 10.1097/CORR.0000000000003185. Epub 2024 Jul 23.

Are Current Survival Prediction Tools Useful When Treating Subsequent Skeletal-related Events From Bone Metastases?当前的生存预测工具在治疗骨转移后的骨骼相关事件时有用吗？

Clin Orthop Relat Res. 2024 Sep 1;482(9):1710-1721. doi: 10.1097/CORR.0000000000003030. Epub 2024 Mar 22.

Signs and symptoms to determine if a patient presenting in primary care or hospital outpatient settings has COVID-19.在基层医疗机构或医院门诊环境中，如果患者出现以下症状和体征，可判断其是否患有 COVID-19。

Cochrane Database Syst Rev. 2022 May 20;5(5):CD013665. doi: 10.1002/14651858.CD013665.pub3.

Cost-effectiveness of using prognostic information to select women with breast cancer for adjuvant systemic therapy.利用预后信息为乳腺癌患者选择辅助性全身治疗的成本效益

Health Technol Assess. 2006 Sep;10(34):iii-iv, ix-xi, 1-204. doi: 10.3310/hta10340.

Rehabilitation following surgery for lumbar spinal stenosis.腰椎管狭窄症手术后的康复

Cochrane Database Syst Rev. 2013 Dec 9;2013(12):CD009644. doi: 10.1002/14651858.CD009644.pub2.

AI-based Hepatic Steatosis Detection and Integrated Hepatic Assessment from Cardiac CT Attenuation Scans Enhances All-cause Mortality Risk Stratification: A Multi-center Study.基于人工智能的心脏CT衰减扫描检测肝脂肪变性及综合肝脏评估可增强全因死亡风险分层：一项多中心研究

medRxiv. 2025 Jun 11:2025.06.09.25329157. doi: 10.1101/2025.06.09.25329157.

Systemic pharmacological treatments for chronic plaque psoriasis: a network meta-analysis.系统性药理学治疗慢性斑块状银屑病：网络荟萃分析。

Cochrane Database Syst Rev. 2021 Apr 19;4(4):CD011535. doi: 10.1002/14651858.CD011535.pub4.

[Volume and health outcomes: evidence from systematic reviews and from evaluation of Italian hospital data].[容量与健康结果：来自系统评价和意大利医院数据评估的证据]

Epidemiol Prev. 2013 Mar-Jun;37(2-3 Suppl 2):1-100.

Comparing the performance of screening surveys versus predictive models in identifying patients in need of health-related social need services in the emergency department.比较筛查调查与预测模型在识别急诊科需要健康相关社会需求服务的患者方面的表现。

PLoS One. 2024 Nov 20;19(11):e0312193. doi: 10.1371/journal.pone.0312193. eCollection 2024.

本文引用的文献

Supportive Care: The "Keystone" of Modern Oncology Practice.支持性护理：现代肿瘤学实践的“基石”。

Cancers (Basel). 2023 Jul 29;15(15):3860. doi: 10.3390/cancers15153860.

Long-term Effect of Machine Learning-Triggered Behavioral Nudges on Serious Illness Conversations and End-of-Life Outcomes Among Patients With Cancer: A Randomized Clinical Trial.机器学习触发的行为推动对癌症患者严重疾病对话和临终结局的长期影响：一项随机临床试验。

JAMA Oncol. 2023 Mar 1;9(3):414-418. doi: 10.1001/jamaoncol.2022.6303.

Two-Phase Sampling Designs for Data Validation in Settings with Covariate Measurement Error and Continuous Outcome.具有协变量测量误差和连续结果的情况下用于数据验证的两阶段抽样设计

J R Stat Soc Ser A Stat Soc. 2021 Oct;184(4):1368-1389. doi: 10.1111/rssa.12689. Epub 2021 Apr 15.

Two-phase stratified sampling and analysis for predicting binary outcomes.两阶段分层抽样分析用于预测二项结局。

Biostatistics. 2023 Jul 14;24(3):585-602. doi: 10.1093/biostatistics/kxab044.

Generalized case-control sampling under generalized linear models.广义线性模型下的广义病例对照抽样。

Biometrics. 2023 Mar;79(1):332-343. doi: 10.1111/biom.13571. Epub 2021 Oct 12.

Optimal Designs of Two-Phase Studies.两阶段研究的最优设计

J Am Stat Assoc. 2020;115(532):1946-1959. doi: 10.1080/01621459.2019.1671200. Epub 2019 Oct 29.

Adoption of clinical risk prediction tools is limited by a lack of integration with electronic health records.临床风险预测工具的应用因缺乏与电子健康记录的整合而受到限制。

BMJ Health Care Inform. 2021 Feb;28(1). doi: 10.1136/bmjhci-2020-100253.

Effect of Integrating Machine Learning Mortality Estimates With Behavioral Nudges to Clinicians on Serious Illness Conversations Among Patients With Cancer: A Stepped-Wedge Cluster Randomized Clinical Trial.将机器学习死亡率估计与行为提示相结合，为临床医生提供指导，以改善癌症患者的严重疾病沟通：一项 stepped-wedge 聚类随机临床试验。

JAMA Oncol. 2020 Dec 1;6(12):e204759. doi: 10.1001/jamaoncol.2020.4759. Epub 2020 Dec 10.

Review of response rates over time in registry-based studies using patient-reported outcome measures.基于患者报告结局测量的注册研究中反应率随时间变化的回顾。

BMJ Open. 2020 Aug 6;10(8):e030808. doi: 10.1136/bmjopen-2019-030808.

Evaluating Discrimination of a Lung Cancer Risk Prediction Model Using Partial Risk-Score in a Two-Phase Study.两阶段研究中使用部分风险评分评估肺癌风险预测模型的判别能力。

Cancer Epidemiol Biomarkers Prev. 2020 Jun;29(6):1196-1203. doi: 10.1158/1055-9965.EPI-19-1574. Epub 2020 Apr 10.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验