从临床记录中提取家族病史信息：深度学习与启发式方法。

Extraction of Family History Information From Clinical Notes: Deep Learning and Heuristics Approach.

作者信息

Silva João Figueira, Almeida João Rafael, Matos Sérgio

机构信息

Department of Electronics, Telecommunications and Informatics, Institute of Electronics and Informatics Engineering of Aveiro, University of Aveiro, Aveiro, Portugal.

Department of Information and Communications Technologies, University of A Coruña, A Coruña, Spain.

出版信息

JMIR Med Inform. 2020 Dec 29;8(12):e22898. doi: 10.2196/22898.

DOI:10.2196/22898

PMID:33372893

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC7803476/

Abstract

BACKGROUND

Electronic health records store large amounts of patient clinical data. Despite efforts to structure patient data, clinical notes containing rich patient information remain stored as free text, greatly limiting its exploitation. This includes family history, which is highly relevant for applications such as diagnosis and prognosis.

OBJECTIVE

This study aims to develop automatic strategies for annotating family history information in clinical notes, focusing not only on the extraction of relevant entities such as family members and disease mentions but also on the extraction of relations between the identified entities.

METHODS

This study extends a previous contribution for the 2019 track on family history extraction from national natural language processing clinical challenges by improving a previously developed rule-based engine, using deep learning (DL) approaches for the extraction of entities from clinical notes, and combining both approaches in a hybrid end-to-end system capable of successfully extracting family member and observation entities and the relations between those entities. Furthermore, this study analyzes the impact of factors such as the use of external resources and different types of embeddings in the performance of DL models.

RESULTS

The approaches developed were evaluated in a first task regarding entity extraction and in a second task concerning relation extraction. The proposed DL approach improved observation extraction, obtaining F scores of 0.8688 and 0.7907 in the training and test sets, respectively. However, DL approaches have limitations in the extraction of family members. The rule-based engine was adjusted to have higher generalizing capability and achieved family member extraction F scores of 0.8823 and 0.8092 in the training and test sets, respectively. The resulting hybrid system obtained F scores of 0.8743 and 0.7979 in the training and test sets, respectively. For the second task, the original evaluator was adjusted to perform a more exact evaluation than the original one, and the hybrid system obtained F scores of 0.6480 and 0.5082 in the training and test sets, respectively.

CONCLUSIONS

We evaluated the impact of several factors on the performance of DL models, and we present an end-to-end system for extracting family history information from clinical notes, which can help in the structuring and reuse of this type of information. The final hybrid solution is provided in a publicly available code repository.

摘要

背景

电子健康记录存储了大量患者临床数据。尽管人们努力对患者数据进行结构化处理，但包含丰富患者信息的临床记录仍以自由文本形式存储，这极大地限制了其利用。这其中包括家族史，而家族史对于诊断和预后等应用非常重要。

目的

本研究旨在开发用于标注临床记录中家族史信息的自动策略，不仅关注家庭成员和疾病提及等相关实体的提取，还关注已识别实体之间关系的提取。

方法

本研究扩展了之前在2019年全国自然语言处理临床挑战中关于家族史提取赛道的贡献，改进了先前开发的基于规则的引擎，使用深度学习（DL）方法从临床记录中提取实体，并将这两种方法结合在一个能够成功提取家庭成员和观察实体以及这些实体之间关系的混合端到端系统中。此外，本研究分析了外部资源的使用和不同类型嵌入等因素对DL模型性能的影响。

结果

所开发的方法在关于实体提取的第一个任务和关于关系提取的第二个任务中进行了评估。所提出的DL方法改进了观察提取，在训练集和测试集中分别获得了0.8688和0.7907的F分数。然而，DL方法在家庭成员提取方面存在局限性。基于规则的引擎经过调整以具有更高的泛化能力，在训练集和测试集中分别实现了0.8823和0.8092的家庭成员提取F分数。最终的混合系统在训练集和测试集中分别获得了0.8743和0.7979的F分数。对于第二个任务，对原始评估器进行了调整，以进行比原始评估更精确的评估，混合系统在训练集和测试集中分别获得了0.6480和0.5082的F分数。

结论

我们评估了几个因素对DL模型性能的影响，并提出了一个从临床记录中提取家族史信息的端到端系统，这有助于此类信息的结构化和重用。最终的混合解决方案在一个公开可用的代码库中提供。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/94a7/7803476/71bfa0cd3b05/medinform_v8i12e22898_fig1.jpg

相似文献

Extraction of Family History Information From Clinical Notes: Deep Learning and Heuristics Approach.从临床记录中提取家族病史信息：深度学习与启发式方法。

JMIR Med Inform. 2020 Dec 29;8(12):e22898. doi: 10.2196/22898.

A Hybrid Model for Family History Information Identification and Relation Extraction: Development and Evaluation of an End-to-End Information Extraction System.一种用于家族病史信息识别与关系抽取的混合模型：一个端到端信息抽取系统的开发与评估

JMIR Med Inform. 2021 Apr 22;9(4):e22797. doi: 10.2196/22797.

Extraction of Information Related to Drug Safety Surveillance From Electronic Health Record Notes: Joint Modeling of Entities and Relations Using Knowledge-Aware Neural Attentive Models.从电子健康记录笔记中提取与药物安全监测相关的信息：使用知识感知神经注意力模型对实体和关系进行联合建模

JMIR Med Inform. 2020 Jul 10;8(7):e18417. doi: 10.2196/18417.

Extraction of Information Related to Adverse Drug Events from Electronic Health Record Notes: Design of an End-to-End Model Based on Deep Learning.从电子健康记录笔记中提取与药物不良事件相关的信息：基于深度学习的端到端模型设计

JMIR Med Inform. 2018 Nov 26;6(4):e12159. doi: 10.2196/12159.

A comparison of word embeddings for the biomedical natural language processing.生物医学自然语言处理中词嵌入的比较。

J Biomed Inform. 2018 Nov;87:12-20. doi: 10.1016/j.jbi.2018.09.008. Epub 2018 Sep 12.

Family History Extraction From Synthetic Clinical Narratives Using Natural Language Processing: Overview and Evaluation of a Challenge Data Set and Solutions for the 2019 National NLP Clinical Challenges (n2c2)/Open Health Natural Language Processing (OHNLP) Competition.利用自然语言处理从合成临床叙述中提取家族病史：2019年国家自然语言处理临床挑战（n2c2）/开放健康自然语言处理（OHNLP）竞赛的挑战数据集概述与评估及解决方案

JMIR Med Inform. 2021 Jan 27;9(1):e24008. doi: 10.2196/24008.

Extracting Drug Names and Associated Attributes From Discharge Summaries: Text Mining Study.从出院小结中提取药物名称及相关属性：文本挖掘研究

JMIR Med Inform. 2021 May 5;9(5):e24678. doi: 10.2196/24678.

Extracting comprehensive clinical information for breast cancer using deep learning methods.利用深度学习方法提取乳腺癌全面临床信息。

Int J Med Inform. 2019 Dec;132:103985. doi: 10.1016/j.ijmedinf.2019.103985. Epub 2019 Oct 2.

Looking for low vision: Predicting visual prognosis by fusing structured and free-text data from electronic health records.寻找低视力：通过融合电子健康记录中的结构化和自由文本数据来预测视觉预后。

Int J Med Inform. 2022 Mar;159:104678. doi: 10.1016/j.ijmedinf.2021.104678. Epub 2021 Dec 30.

Clinical Relation Extraction Toward Drug Safety Surveillance Using Electronic Health Record Narratives: Classical Learning Versus Deep Learning.利用电子健康记录叙述进行药物安全监测的临床关系提取：经典学习与深度学习

JMIR Public Health Surveill. 2018 Apr 25;4(2):e29. doi: 10.2196/publichealth.9361.

引用本文的文献

A comparison of family health history of breast cancer, colorectal cancer, and diabetes in self-reported survey and electronic health records data, All of Us Research Program.“我们所有人”研究项目中自我报告调查与电子健康记录数据里乳腺癌、结直肠癌和糖尿病家族健康史的比较

Genet Med Open. 2025 Jun 11;3:103439. doi: 10.1016/j.gimo.2025.103439. eCollection 2025.

Unsupervised SapBERT-based bi-encoders for medical concept annotation of clinical narratives with SNOMED CT.基于无监督SapBERT的双编码器，用于使用SNOMED CT对临床叙述进行医学概念注释。

Digit Health. 2024 Oct 21;10:20552076241288681. doi: 10.1177/20552076241288681. eCollection 2024 Jan-Dec.

Health Care Language Models and Their Fine-Tuning for Information Extraction: Scoping Review.医疗保健语言模型及其在信息提取方面的微调：范围综述。

JMIR Med Inform. 2024 Oct 21;12:e60164. doi: 10.2196/60164.

Automated Family Histories Significantly Improve Risk Prediction in an EHR.自动化家族病史显著改善电子健康记录中的风险预测。

AMIA Jt Summits Transl Sci Proc. 2024 May 31;2024:221-229. eCollection 2024.

Developing a Natural Language Processing tool to identify perinatal self-harm in electronic healthcare records.开发一种自然语言处理工具，以在电子医疗记录中识别围产期自伤。

PLoS One. 2021 Aug 4;16(8):e0253809. doi: 10.1371/journal.pone.0253809. eCollection 2021.

本文引用的文献

Clinical Text Data in Machine Learning: Systematic Review.机器学习中的临床文本数据：系统综述

JMIR Med Inform. 2020 Mar 31;8(3):e17984. doi: 10.2196/17984.

Family history information extraction via deep joint learning.通过深度联合学习提取家族史信息。

BMC Med Inform Decis Mak. 2019 Dec 27;19(Suppl 10):277. doi: 10.1186/s12911-019-0995-5.

An attention-based deep learning model for clinical named entity recognition of Chinese electronic medical records.基于注意力的深度学习模型在中文电子病历临床命名实体识别中的应用。

BMC Med Inform Decis Mak. 2019 Dec 5;19(Suppl 5):235. doi: 10.1186/s12911-019-0933-6.

BioBERT: a pre-trained biomedical language representation model for biomedical text mining.BioBERT：一种用于生物医学文本挖掘的预训练生物医学语言表示模型。

Bioinformatics. 2020 Feb 15;36(4):1234-1240. doi: 10.1093/bioinformatics/btz682.

BioWordVec, improving biomedical word embeddings with subword information and MeSH.BioWordVec，利用子词信息和 MeSH 改进生物医学词向量。

Sci Data. 2019 May 10;6(1):52. doi: 10.1038/s41597-019-0055-0.

Natural Language Processing of Clinical Notes on Chronic Diseases: Systematic Review.慢性病临床记录的自然语言处理：系统综述

JMIR Med Inform. 2019 Apr 27;7(2):e12239. doi: 10.2196/12239.

Configurable web-services for biomedical document annotation.用于生物医学文档注释的可配置网络服务。

J Cheminform. 2018 Dec 21;10(1):68. doi: 10.1186/s13321-018-0317-4.

Opportunities and obstacles for deep learning in biology and medicine.深度学习在生物学和医学中的机遇与挑战。

J R Soc Interface. 2018 Apr;15(141). doi: 10.1098/rsif.2017.0387.

CLAMP - a toolkit for efficiently building customized clinical natural language processing pipelines.CLAMP - 一个用于高效构建定制化临床自然语言处理管道的工具包。

J Am Med Inform Assoc. 2018 Mar 1;25(3):331-336. doi: 10.1093/jamia/ocx132.

Systematic Analysis of Free-Text Family History in Electronic Health Record.电子健康记录中自由文本家族病史的系统分析

AMIA Jt Summits Transl Sci Proc. 2017 Jul 26;2017:104-113. eCollection 2017.

文献AI研究员

20分钟写一篇综述，助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型，支持多种主流文档格式。

立即体验

从临床记录中提取家族病史信息：深度学习与启发式方法。

Extraction of Family History Information From Clinical Notes: Deep Learning and Heuristics Approach.

作者信息

机构信息

出版信息

BACKGROUND

OBJECTIVE

METHODS

RESULTS

CONCLUSIONS

背景

目的

方法

结果

结论

相似文献

引用本文的文献

本文引用的文献

文献AI研究员

用中文搜PubMed

文档翻译

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献