利用组合对预测异源二聚体蛋白复合物进行改进，使用成对核函数。

Improving prediction of heterodimeric protein complexes using combination with pairwise kernel.

机构信息

Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), Tokyo, Japan.

Department of Electrical Engineering and Computer Science, National Institute of Technology, Matsue College, 14-4, Nishiikumacho, Matsue, 690-8518, Japan.

出版信息

BMC Bioinformatics. 2018 Feb 19;19(Suppl 1):39. doi: 10.1186/s12859-018-2017-5.

DOI:10.1186/s12859-018-2017-5

PMID:29504897

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC5836830/

Abstract

BACKGROUND

Since many proteins become functional only after they interact with their partner proteins and form protein complexes, it is essential to identify the sets of proteins that form complexes. Therefore, several computational methods have been proposed to predict complexes from the topology and structure of experimental protein-protein interaction (PPI) network. These methods work well to predict complexes involving at least three proteins, but generally fail at identifying complexes involving only two different proteins, called heterodimeric complexes or heterodimers. There is however an urgent need for efficient methods to predict heterodimers, since the majority of known protein complexes are precisely heterodimers.

RESULTS

In this paper, we use three promising kernel functions, Min kernel and two pairwise kernels, which are Metric Learning Pairwise Kernel (MLPK) and Tensor Product Pairwise Kernel (TPPK). We also consider the normalization forms of Min kernel. Then, we combine Min kernel or its normalization form and one of the pairwise kernels by plugging. We applied kernels based on PPI, domain, phylogenetic profile, and subcellular localization properties to predicting heterodimers. Then, we evaluate our method by employing C-Support Vector Classification (C-SVC), carrying out 10-fold cross-validation, and calculating the average F-measures. The results suggest that the combination of normalized-Min-kernel and MLPK leads to the best F-measure and improved the performance of our previous work, which had been the best existing method so far.

CONCLUSIONS

We propose new methods to predict heterodimers, using a machine learning-based approach. We train a support vector machine (SVM) to discriminate interacting vs non-interacting protein pairs, based on informations extracted from PPI, domain, phylogenetic profiles and subcellular localization. We evaluate in detail new kernel functions to encode these data, and report prediction performance that outperforms the state-of-the-art.

摘要

背景

由于许多蛋白质只有在与它们的伴侣蛋白质相互作用并形成蛋白质复合物后才具有功能，因此识别形成复合物的蛋白质组是至关重要的。因此，已经提出了几种计算方法来根据实验蛋白质-蛋白质相互作用（PPI）网络的拓扑结构和结构来预测复合物。这些方法在预测涉及至少三种蛋白质的复合物时效果很好，但通常无法识别仅涉及两种不同蛋白质的复合物，称为异二聚体复合物或异二聚体。然而，迫切需要有效的方法来预测异二聚体，因为大多数已知的蛋白质复合物正是异二聚体。

结果

在本文中，我们使用了三种有前途的核函数，Min 核和两种成对核函数，即度量学习成对核函数（MLPK）和张量积成对核函数（TPPK）。我们还考虑了 Min 核的归一化形式。然后，我们通过插入将 Min 核或其归一化形式与一个成对核函数组合在一起。我们基于 PPI、结构域、系统发生谱和亚细胞定位特性应用核函数来预测异二聚体。然后，我们通过使用 C-支持向量分类器（C-SVC）、进行 10 折交叉验证并计算平均 F 度量来评估我们的方法。结果表明，归一化-Min 核与 MLPK 的组合导致最佳 F 度量，并提高了我们之前的工作的性能，这是迄今为止最好的现有方法。

结论

我们提出了使用基于机器学习的方法预测异二聚体的新方法。我们基于从 PPI、结构域、系统发生谱和亚细胞定位中提取的信息，使用支持向量机（SVM）来区分相互作用和非相互作用的蛋白质对。我们详细评估了新的核函数来编码这些数据，并报告了优于最先进方法的预测性能。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/b1b5/5836830/c277d07441f1/12859_2018_2017_Fig1_HTML.jpg

相似文献

Improving prediction of heterodimeric protein complexes using combination with pairwise kernel.利用组合对预测异源二聚体蛋白复合物进行改进，使用成对核函数。

BMC Bioinformatics. 2018 Feb 19;19(Suppl 1):39. doi: 10.1186/s12859-018-2017-5.

Prediction of heterotrimeric protein complexes by two-phase learning using neighboring kernels.利用邻域核的两阶段学习预测异源三聚体蛋白复合物。

BMC Bioinformatics. 2014;15 Suppl 2(Suppl 2):S6. doi: 10.1186/1471-2105-15-S2-S6. Epub 2014 Jan 24.

Metabolic network prediction through pairwise rational kernels.通过成对有理核进行代谢网络预测。

BMC Bioinformatics. 2014 Sep 26;15(1):318. doi: 10.1186/1471-2105-15-318.

Prediction of heterodimeric protein complexes from weighted protein-protein interaction networks using novel features and kernel functions.基于新型特征和核函数的加权蛋白质-蛋白质相互作用网络预测异源二聚体蛋白复合物。

PLoS One. 2013 Jun 11;8(6):e65265. doi: 10.1371/journal.pone.0065265. Print 2013.

Characterizing informative sequence descriptors and predicting binding affinities of heterodimeric protein complexes.表征信息性序列描述符并预测异源二聚体蛋白复合物的结合亲和力。

BMC Bioinformatics. 2015;16 Suppl 18(Suppl 18):S14. doi: 10.1186/1471-2105-16-S18-S14. Epub 2015 Dec 9.

Prediction of protein-protein interaction with pairwise kernel support vector machine.基于成对核支持向量机的蛋白质-蛋白质相互作用预测

Int J Mol Sci. 2014 Feb 21;15(2):3220-33. doi: 10.3390/ijms15023220.

A new pairwise kernel for biological network inference with support vector machines.一种用于支持向量机生物网络推理的新型成对核。

BMC Bioinformatics. 2007;8 Suppl 10(Suppl 10):S8. doi: 10.1186/1471-2105-8-S10-S8.

Protein subcellular localization prediction using multiple kernel learning based support vector machine.基于多核学习支持向量机的蛋白质亚细胞定位预测

Mol Biosyst. 2017 Mar 28;13(4):785-795. doi: 10.1039/c6mb00860g.

Feature selection and classification of protein-protein complexes based on their binding affinities using machine learning approaches.基于机器学习方法，利用蛋白质-蛋白质复合物的结合亲和力进行特征选择和分类。

Proteins. 2014 Sep;82(9):2088-96. doi: 10.1002/prot.24564. Epub 2014 Apr 16.

Predicting co-complexed protein pairs from heterogeneous data.从异构数据中预测共复合蛋白质对。

PLoS Comput Biol. 2008 Apr 18;4(4):e1000054. doi: 10.1371/journal.pcbi.1000054.

引用本文的文献

Unravelling the human taste receptor interactome: machine learning and molecular modelling insights into protein-protein interactions.解析人类味觉受体相互作用组：机器学习与蛋白质-蛋白质相互作用的分子建模见解

NPJ Sci Food. 2025 Jul 1;9(1):113. doi: 10.1038/s41538-025-00478-9.

Identification of all-against-all protein-protein interactions based on deep hash learning.基于深度哈希学习的全对全蛋白质-蛋白质相互作用识别。

BMC Bioinformatics. 2022 Jul 8;23(1):266. doi: 10.1186/s12859-022-04811-x.

Efficient and accurate identification of protein complexes from protein-protein interaction networks based on the clustering coefficient.基于聚类系数从蛋白质-蛋白质相互作用网络中高效准确地识别蛋白质复合物。

Comput Struct Biotechnol J. 2021 Sep 20;19:5255-5263. doi: 10.1016/j.csbj.2021.09.014. eCollection 2021.

Detecting overlapping protein complexes in weighted PPI network based on overlay network chain in quotient space.基于商空间重叠网络链的加权 PPI 网络中重叠蛋白质复合物的检测。

BMC Bioinformatics. 2019 Dec 24;20(Suppl 25):682. doi: 10.1186/s12859-019-3256-9.

Machine-learning techniques for the prediction of protein-protein interactions.基于机器学习的蛋白质-蛋白质相互作用预测技术。

J Biosci. 2019 Sep;44(4).

本文引用的文献

Discovery of small protein complexes from PPI networks with size-specific supervised weighting.通过大小特异性监督加权从蛋白质-蛋白质相互作用网络中发现小蛋白质复合物。

BMC Syst Biol. 2014;8 Suppl 5(Suppl 5):S3. doi: 10.1186/1752-0509-8-S5-S3. Epub 2014 Dec 12.

Proteins. 2014 Sep;82(9):2088-96. doi: 10.1002/prot.24564. Epub 2014 Apr 16.

PLoS One. 2013 Jun 11;8(6):e65265. doi: 10.1371/journal.pone.0065265. Print 2013.

KEGG OC: a large-scale automatic construction of taxonomy-based ortholog clusters.KEGG Orthology（KO）：一个大规模的基于分类学的直系同源簇自动构建方法。

Nucleic Acids Res. 2013 Jan;41(Database issue):D353-7. doi: 10.1093/nar/gks1239. Epub 2012 Nov 27.

NWE: Node-weighted expansion for protein complex prediction using random walk distances.NWE：基于随机游走距离的节点加权扩展的蛋白质复合物预测方法。

Proteome Sci. 2011 Oct 14;9 Suppl 1(Suppl 1):S14. doi: 10.1186/1477-5956-9-S1-S14.

DIPOS: database of interacting proteins in Oryza sativa.DIPOS：水稻相互作用蛋白质数据库。

Mol Biosyst. 2011 Sep;7(9):2615-21. doi: 10.1039/c1mb05120b. Epub 2011 Jun 29.

A max-flow-based approach to the identification of protein complexes using protein interaction and microarray data.基于最大流的方法，利用蛋白质相互作用和微阵列数据鉴定蛋白质复合物。

IEEE/ACM Trans Comput Biol Bioinform. 2011 May-Jun;8(3):621-34. doi: 10.1109/TCBB.2010.78.

RRW: repeated random walks on genome-scale protein networks for local cluster discovery.RRW：基于全基因组尺度蛋白质网络的重复随机游走用于局部簇发现。

BMC Bioinformatics. 2009 Sep 9;10:283. doi: 10.1186/1471-2105-10-283.

FPPI: Fusarium graminearum protein-protein interaction database.FPPI：禾谷镰刀菌蛋白-蛋白相互作用数据库。

J Proteome Res. 2009 Oct;8(10):4714-21. doi: 10.1021/pr900415b.

Complex discovery from weighted PPI networks.基于加权 PPI 网络的复杂发现。

Bioinformatics. 2009 Aug 1;25(15):1891-7. doi: 10.1093/bioinformatics/btp311. Epub 2009 May 12.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

利用组合对预测异源二聚体蛋白复合物进行改进，使用成对核函数。

Improving prediction of heterodimeric protein complexes using combination with pairwise kernel.

机构信息

出版信息

BACKGROUND

RESULTS

CONCLUSIONS

背景

结果

结论

相似文献

引用本文的文献

本文引用的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献