利用一种新颖的迭代自适应稀疏偏最小二乘算法识别原核基因组的短编码序列。

Recognizing short coding sequences of prokaryotic genome using a novel iteratively adaptive sparse partial least squares algorithm.

机构信息

School of Chemical Engineering and Technology, Tianjin University, Tianjin 300072, China.

出版信息

Biol Direct. 2013 Sep 25;8:23. doi: 10.1186/1745-6150-8-23.

DOI:10.1186/1745-6150-8-23

PMID:24067167

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC3852556/

Abstract

BACKGROUND

Significant efforts have been made to address the problem of identifying short genes in prokaryotic genomes. However, most known methods are not effective in detecting short genes. Because of the limited information contained in short DNA sequences, it is very difficult to accurately distinguish between protein coding and non-coding sequences in prokaryotic genomes. We have developed a new Iteratively Adaptive Sparse Partial Least Squares (IASPLS) algorithm as the classifier to improve the accuracy of the identification process.

RESULTS

For testing, we chose the short coding and non-coding sequences from seven prokaryotic organisms. We used seven feature sets (including GC content, Z-curve, etc.) of short genes.In comparison with GeneMarkS, Metagene, Orphelia, and Heuristic Approachs methods, our model achieved the best prediction performance in identification of short prokaryotic genes. Even when we focused on the very short length group ([60-100 nt)), our model provided sensitivity as high as 83.44% and specificity as high as 92.8%. These values are two or three times higher than three of the other methods while Metagene fails to recognize genes in this length range.The experiments also proved that the IASPLS can improve the identification accuracy in comparison with other widely used classifiers, i.e. Logistic, Random Forest (RF) and K nearest neighbors (KNN). The accuracy in using IASPLS was improved 5.90% or more in comparison with the other methods. In addition to the improvements in accuracy, IASPLS required ten times less computer time than using KNN or RF.

CONCLUSIONS

It is conclusive that our method is preferable for application as an automated method of short gene classification. Its linearity and easily optimized parameters make it practicable for predicting short genes of newly-sequenced or under-studied species.

摘要

背景

在识别原核基因组中的短基因方面已经做出了巨大努力。然而，大多数已知的方法在检测短基因方面并不有效。由于短 DNA 序列中包含的信息量有限，因此很难准确区分原核基因组中的蛋白质编码序列和非编码序列。我们开发了一种新的迭代自适应稀疏偏最小二乘（IASPLS）算法作为分类器，以提高识别过程的准确性。

结果

为了测试，我们从七个原核生物中选择了短编码和非编码序列。我们使用了七个短基因特征集（包括 GC 含量、Z 曲线等）。与 GeneMarkS、Metagene、Orphelia 和启发式方法相比，我们的模型在短原核基因的识别中表现出了最佳的预测性能。即使我们专注于非常短的长度组（[60-100 nt）），我们的模型提供的灵敏度也高达 83.44%，特异性高达 92.8%。这些值比其他三种方法高出两到三倍，而 Metagene 无法识别这个长度范围内的基因。实验还证明，与其他广泛使用的分类器（即 Logistic、随机森林（RF）和 K 最近邻（KNN））相比，IASPLS 可以提高识别精度。与其他方法相比，使用 IASPLS 的精度提高了 5.90%以上。除了准确性的提高之外，IASPLS 所需的计算机时间比使用 KNN 或 RF 少十倍。

结论

我们的方法可作为短基因分类的自动化方法，这是一个明确的结论。其线性和易于优化的参数使其适用于预测新测序或研究较少的物种的短基因。

相似文献

Recognizing short coding sequences of prokaryotic genome using a novel iteratively adaptive sparse partial least squares algorithm.利用一种新颖的迭代自适应稀疏偏最小二乘算法识别原核基因组的短编码序列。

Biol Direct. 2013 Sep 25;8:23. doi: 10.1186/1745-6150-8-23.

Classifier assessment and feature selection for recognizing short coding sequences of human genes.用于识别人类基因短编码序列的分类器评估与特征选择

J Comput Biol. 2012 Mar;19(3):251-60. doi: 10.1089/cmb.2011.0078.

Finding prokaryotic genes by the 'frame-by-frame' algorithm: targeting gene starts and overlapping genes.通过“逐帧”算法寻找原核生物基因：靶向基因起始位点和重叠基因。

Bioinformatics. 1999 Nov;15(11):874-86. doi: 10.1093/bioinformatics/15.11.874.

Predicting essential genes in prokaryotic genomes using a linear method: ZUPLS.使用线性方法ZUPLS预测原核生物基因组中的必需基因。

Integr Biol (Camb). 2014 Apr;6(4):460-9. doi: 10.1039/c3ib40241j. Epub 2014 Mar 7.

[Comprehensive re-annotation of protein-coding genes for prokaryotic genomes by Z-curve and similarity-based methods].[基于Z曲线和相似性方法对原核生物基因组蛋白质编码基因进行全面重新注释]

Yi Chuan. 2020 Jul 20;42(7):691-702. doi: 10.16288/j.yczz.20-022.

A Partial Least Squares Based Procedure for Upstream Sequence Classification in Prokaryotes.

IEEE/ACM Trans Comput Biol Bioinform. 2015 May-Jun;12(3):560-7. doi: 10.1109/TCBB.2014.2366146.

Multivariate entropy distance method for prokaryotic gene identification.用于原核基因识别的多变量熵距离方法

J Bioinform Comput Biol. 2004 Jun;2(2):353-73. doi: 10.1142/s0219720004000624.

Comparison of various algorithms for recognizing short coding sequences of human genes.用于识别人类基因短编码序列的各种算法的比较。

Bioinformatics. 2004 Mar 22;20(5):673-81. doi: 10.1093/bioinformatics/btg467. Epub 2004 Feb 5.

The elusive short gene--an ensemble method for recognition for prokaryotic genome. elusive 短基因——原核基因组识别的集成方法。

Biochem Biophys Res Commun. 2012 May 25;422(1):36-41. doi: 10.1016/j.bbrc.2012.04.090. Epub 2012 Apr 25.

引用本文的文献

Alternative ORFs and small ORFs: shedding light on the dark proteome.替代开放阅读框和小开放阅读框：揭示暗蛋白质组的奥秘。

Nucleic Acids Res. 2020 Feb 20;48(3):1029-1042. doi: 10.1093/nar/gkz734.

Gene Prediction in Metagenomic Fragments with Deep Learning.利用深度学习进行宏基因组片段中的基因预测

Biomed Res Int. 2017;2017:4740354. doi: 10.1155/2017/4740354. Epub 2017 Nov 8.

Comparative genomic analysis shows that avian pathogenic Escherichia coli isolate IMT5155 (O2:K1:H5; ST complex 95, ST140) shares close relationship with ST95 APEC O1:K1 and human ExPEC O18:K1 strains.比较基因组分析表明，禽致病性大肠杆菌分离株IMT5155（O2:K1:H5；ST复合体95，ST140）与ST95 APEC O1:K1和人源肠外致病性大肠杆菌O18:K1菌株关系密切。

PLoS One. 2014 Nov 14;9(11):e112048. doi: 10.1371/journal.pone.0112048. eCollection 2014.

本文引用的文献

The elusive short gene--an ensemble method for recognition for prokaryotic genome. elusive 短基因——原核基因组识别的集成方法。

Biochem Biophys Res Commun. 2012 May 25;422(1):36-41. doi: 10.1016/j.bbrc.2012.04.090. Epub 2012 Apr 25.

Classifier assessment and feature selection for recognizing short coding sequences of human genes.用于识别人类基因短编码序列的分类器评估与特征选择

J Comput Biol. 2012 Mar;19(3):251-60. doi: 10.1089/cmb.2011.0078.

Recognition of prokaryotic promoters based on a novel variable-window Z-curve method.基于新型可变窗口 Z 曲线方法的原核启动子识别。

Nucleic Acids Res. 2012 Feb;40(3):963-71. doi: 10.1093/nar/gkr795. Epub 2011 Sep 27.

Identification of prokaryotic small proteins using a comparative genomic approach.利用比较基因组学方法鉴定原核生物小蛋白。

Bioinformatics. 2011 Jul 1;27(13):1765-71. doi: 10.1093/bioinformatics/btr275. Epub 2011 May 5.

Ab initio gene identification in metagenomic sequences.从头鉴定宏基因组序列中的基因。

Nucleic Acids Res. 2010 Jul;38(12):e132. doi: 10.1093/nar/gkq275. Epub 2010 Apr 19.

Prodigal: prokaryotic gene recognition and translation initiation site identification.普罗迪格：原核基因识别和翻译起始位点鉴定。

BMC Bioinformatics. 2010 Mar 8;11:119. doi: 10.1186/1471-2105-11-119.

Orphelia: predicting genes in metagenomic sequencing reads.奥菲莉亚：宏基因组测序读段中的基因预测

Nucleic Acids Res. 2009 Jul;37(Web Server issue):W101-5. doi: 10.1093/nar/gkp327. Epub 2009 May 8.

Small membrane proteins found by comparative genomics and ribosome binding site models.通过比较基因组学和核糖体结合位点模型发现的小膜蛋白。

Mol Microbiol. 2008 Dec;70(6):1487-501. doi: 10.1111/j.1365-2958.2008.06495.x.

A sparse PLS for variable selection when integrating omics data.整合组学数据时用于变量选择的稀疏偏最小二乘法

Stat Appl Genet Mol Biol. 2008;7(1):Article 35. doi: 10.2202/1544-6115.1390. Epub 2008 Nov 18.

DiProDB: a database for dinucleotide properties.DiProDB：一个关于二核苷酸特性的数据库。

Nucleic Acids Res. 2009 Jan;37(Database issue):D37-40. doi: 10.1093/nar/gkn597. Epub 2008 Sep 19.

文献AI研究员

20分钟写一篇综述，助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型，支持多种主流文档格式。

立即体验

利用一种新颖的迭代自适应稀疏偏最小二乘算法识别原核基因组的短编码序列。

Recognizing short coding sequences of prokaryotic genome using a novel iteratively adaptive sparse partial least squares algorithm.

机构信息

出版信息

BACKGROUND

RESULTS

CONCLUSIONS

背景

结果

结论

相似文献

引用本文的文献

本文引用的文献

文献AI研究员

用中文搜PubMed

文档翻译

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献