一种用于改进基于深度学习的定量构效关系回归建模中不确定性量化的混合框架。

A hybrid framework for improving uncertainty quantification in deep learning-based QSAR regression modeling.

作者信息

Wang Dingyan, Yu Jie, Chen Lifan, Li Xutong, Jiang Hualiang, Chen Kaixian, Zheng Mingyue, Luo Xiaomin

机构信息

Shanghai Key Laboratory of Forensic Medicine, Academy of Forensic Science, Shanghai, 200063, China.

University of Chinese Academy of Sciences, No.19A Yuquan Road, Beijing, 100049, China.

出版信息

J Cheminform. 2021 Sep 20;13(1):69. doi: 10.1186/s13321-021-00551-x.

DOI:10.1186/s13321-021-00551-x

PMID:34544485

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC8454160/

Abstract

Reliable uncertainty quantification for statistical models is crucial in various downstream applications, especially for drug design and discovery where mistakes may incur a large amount of cost. This topic has therefore absorbed much attention and a plethora of methods have been proposed over the past years. The approaches that have been reported so far can be mainly categorized into two classes: distance-based approaches and Bayesian approaches. Although these methods have been widely used in many scenarios and shown promising performance with their distinct superiorities, being overconfident on out-of-distribution examples still poses challenges for the deployment of these techniques in real-world applications. In this study we investigated a number of consensus strategies in order to combine both distance-based and Bayesian approaches together with post-hoc calibration for improved uncertainty quantification in QSAR (Quantitative Structure-Activity Relationship) regression modeling. We employed a set of criteria to quantitatively assess the ranking and calibration ability of these models. Experiments based on 24 bioactivity datasets were designed to make critical comparison between the model we proposed and other well-studied baseline models. Our findings indicate that the hybrid framework proposed by us can robustly enhance the model ability of ranking absolute errors. Together with post-hoc calibration on the validation set, we show that well-calibrated uncertainty quantification results can be obtained in domain shift settings. The complementarity between different methods is also conceptually analyzed.

摘要

对于统计模型而言，可靠的不确定性量化在各种下游应用中至关重要，尤其是在药物设计与发现领域，因为错误可能会导致巨额成本。因此，这一主题备受关注，在过去几年中人们提出了大量方法。目前已报道的方法主要可分为两类：基于距离的方法和贝叶斯方法。尽管这些方法已在许多场景中广泛使用，并凭借其独特优势展现出了良好的性能，但对分布外样本过度自信仍然给这些技术在实际应用中的部署带来了挑战。在本研究中，我们研究了多种共识策略，以便将基于距离的方法和贝叶斯方法与事后校准相结合，从而在定量构效关系（QSAR）回归建模中改进不确定性量化。我们采用了一组标准来定量评估这些模型的排序和校准能力。基于24个生物活性数据集设计了实验，以对我们提出的模型与其他经过充分研究的基线模型进行关键比较。我们的研究结果表明，我们提出的混合框架能够有力地增强模型对绝对误差的排序能力。结合在验证集上的事后校准，我们表明在域转移设置中可以获得校准良好的不确定性量化结果。我们还从概念上分析了不同方法之间的互补性。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/defc/8454160/d39e4ec11839/13321_2021_551_Fig1_HTML.jpg

相似文献

A hybrid framework for improving uncertainty quantification in deep learning-based QSAR regression modeling.一种用于改进基于深度学习的定量构效关系回归建模中不确定性量化的混合框架。

J Cheminform. 2021 Sep 20;13(1):69. doi: 10.1186/s13321-021-00551-x.

Reducing overconfident errors in molecular property classification using Posterior Network.使用后验网络减少分子性质分类中的过度自信错误。

Patterns (N Y). 2024 May 8;5(6):100991. doi: 10.1016/j.patter.2024.100991. eCollection 2024 Jun 14.

Empirical Frequentist Coverage of Deep Learning Uncertainty Quantification Procedures.深度学习不确定性量化程序的经验频率主义覆盖率

Entropy (Basel). 2021 Nov 30;23(12):1608. doi: 10.3390/e23121608.

Uncertainty-based saltwater intrusion prediction using integrated Bayesian machine learning modeling (IBMLM) in a deep aquifer.基于不确定性的深层含水层海水入侵预测：综合贝叶斯机器学习模型（IBMLM）的应用。

J Environ Manage. 2024 Mar;354:120252. doi: 10.1016/j.jenvman.2024.120252. Epub 2024 Feb 22.

Folic acid supplementation and malaria susceptibility and severity among people taking antifolate antimalarial drugs in endemic areas.在流行地区，服用抗叶酸抗疟药物的人群中，叶酸补充剂与疟疾易感性和严重程度的关系。

Cochrane Database Syst Rev. 2022 Feb 1;2(2022):CD014217. doi: 10.1002/14651858.CD014217.

Click-through Rate Prediction and Uncertainty Quantification Based on Bayesian Deep Learning.基于贝叶斯深度学习的点击率预测与不确定性量化

Entropy (Basel). 2023 Feb 23;25(3):406. doi: 10.3390/e25030406.

Leveraging Bayesian deep learning and ensemble methods for uncertainty quantification in image classification: A ranking-based approach.利用贝叶斯深度学习和集成方法进行图像分类中的不确定性量化：一种基于排序的方法。

Heliyon. 2024 Jan 8;10(2):e24188. doi: 10.1016/j.heliyon.2024.e24188. eCollection 2024 Jan 30.

Evaluating Scalable Uncertainty Estimation Methods for Deep Learning-Based Molecular Property Prediction.评估基于深度学习的分子性质预测的可扩展不确定性估计方法。

J Chem Inf Model. 2020 Jun 22;60(6):2697-2717. doi: 10.1021/acs.jcim.9b00975. Epub 2020 Apr 24.

Deep convolutional neural network and IoT technology for healthcare.用于医疗保健的深度卷积神经网络和物联网技术。

Digit Health. 2024 Jan 17;10:20552076231220123. doi: 10.1177/20552076231220123. eCollection 2024 Jan-Dec.

DropConnect is effective in modeling uncertainty of Bayesian deep networks.DropConnect 在对贝叶斯深度网络的不确定性建模方面非常有效。

Sci Rep. 2021 Mar 9;11(1):5458. doi: 10.1038/s41598-021-84854-x.

引用本文的文献

A data-driven generative strategy to avoid reward hacking in multi-objective molecular design.一种数据驱动的生成策略，用于避免多目标分子设计中的奖励操纵。

Nat Commun. 2025 Mar 11;16(1):2409. doi: 10.1038/s41467-025-57582-3.

Achieving well-informed decision-making in drug discovery: a comprehensive calibration study using neural network-based structure-activity models.在药物发现中实现明智的决策：一项使用基于神经网络的构效模型的全面校准研究。

J Cheminform. 2025 Mar 5;17(1):29. doi: 10.1186/s13321-025-00964-y.

Reducing overconfident errors in molecular property classification using Posterior Network.使用后验网络减少分子性质分类中的过度自信错误。

Patterns (N Y). 2024 May 8;5(6):100991. doi: 10.1016/j.patter.2024.100991. eCollection 2024 Jun 14.

An overview of recent advances and challenges in predicting compound-protein interaction (CPI).预测化合物-蛋白质相互作用（CPI）的最新进展与挑战概述。

Med Rev (2021). 2023 Oct 6;3(6):465-486. doi: 10.1515/mr-2023-0030. eCollection 2023 Dec.

Chemprop: A Machine Learning Package for Chemical Property Prediction.Chemprop：一个用于化学性质预测的机器学习工具包。

J Chem Inf Model. 2024 Jan 8;64(1):9-17. doi: 10.1021/acs.jcim.3c01250. Epub 2023 Dec 26.

Recent Deep Learning Applications to Structure-Based Drug Design.基于结构的药物设计的最新深度学习应用。

Methods Mol Biol. 2024;2714:215-234. doi: 10.1007/978-1-0716-3441-7_13.

Evaluating point-prediction uncertainties in neural networks for protein-ligand binding prediction.评估用于蛋白质-配体结合预测的神经网络中的点预测不确定性。

Artif Intell Chem. 2023 Jun;1(1). doi: 10.1016/j.aichem.2023.100004. Epub 2023 Jun 3.

Large-scale evaluation of k-fold cross-validation ensembles for uncertainty estimation.用于不确定性估计的k折交叉验证集成的大规模评估。

J Cheminform. 2023 Apr 28;15(1):49. doi: 10.1186/s13321-023-00709-9.

Convolutional Neural Network Model Based on 2D Fingerprint for Bioactivity Prediction.基于二维指纹的卷积神经网络模型用于生物活性预测。

Int J Mol Sci. 2022 Oct 30;23(21):13230. doi: 10.3390/ijms232113230.

Uncertainty quantification: Can we trust artificial intelligence in drug discovery?不确定性量化：在药物研发中我们能信任人工智能吗？

iScience. 2022 Jul 21;25(8):104814. doi: 10.1016/j.isci.2022.104814. eCollection 2022 Aug 19.

本文引用的文献

Evaluating and Calibrating Uncertainty Prediction in Regression Tasks.评估与校准回归任务中的不确定性预测

Sensors (Basel). 2022 Jul 25;22(15):5540. doi: 10.3390/s22155540.

Assigning confidence to molecular property prediction.为分子性质预测分配置信度。

Expert Opin Drug Discov. 2021 Sep;16(9):1009-1023. doi: 10.1080/17460441.2021.1925247. Epub 2021 Jun 15.

Could graph neural networks learn better molecular representation for drug discovery? A comparison study of descriptor-based and graph-based models.图神经网络能否为药物发现学习更好的分子表示？基于描述符和基于图的模型的比较研究。

J Cheminform. 2021 Feb 17;13(1):12. doi: 10.1186/s13321-020-00479-8.

Predicting materials properties without crystal structure: deep representation learning from stoichiometry.无需晶体结构预测材料属性：基于化学计量学的深度表征学习

Nat Commun. 2020 Dec 8;11(1):6280. doi: 10.1038/s41467-020-19964-7.

Uncertainty quantification in drug design.药物设计中的不确定性量化。

Drug Discov Today. 2021 Feb;26(2):474-489. doi: 10.1016/j.drudis.2020.11.027. Epub 2020 Nov 27.

Leveraging Uncertainty in Machine Learning Accelerates Biological Discovery and Design.利用机器学习中的不确定性加速生物学发现和设计。

Cell Syst. 2020 Nov 18;11(5):461-477.e9. doi: 10.1016/j.cels.2020.09.007. Epub 2020 Oct 15.

Comprehensive Analysis of Applicability Domains of QSPR Models for Chemical Reactions.全面分析 QSPR 模型在化学反应中的适用性域。

Int J Mol Sci. 2020 Aug 3;21(15):5542. doi: 10.3390/ijms21155542.

Uncertainty Quantification Using Neural Networks for Molecular Property Prediction.使用神经网络进行分子性质预测的不确定性量化。

J Chem Inf Model. 2020 Aug 24;60(8):3770-3780. doi: 10.1021/acs.jcim.0c00502. Epub 2020 Aug 4.

QSAR without borders.无边界定量构效关系。

Chem Soc Rev. 2020 Jun 7;49(11):3525-3564. doi: 10.1039/d0cs00098a. Epub 2020 May 1.

Evaluating Scalable Uncertainty Estimation Methods for Deep Learning-Based Molecular Property Prediction.评估基于深度学习的分子性质预测的可扩展不确定性估计方法。

J Chem Inf Model. 2020 Jun 22;60(6):2697-2717. doi: 10.1021/acs.jcim.9b00975. Epub 2020 Apr 24.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

一种用于改进基于深度学习的定量构效关系回归建模中不确定性量化的混合框架。

A hybrid framework for improving uncertainty quantification in deep learning-based QSAR regression modeling.

作者信息

机构信息

出版信息

相似文献

引用本文的文献

本文引用的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献