Suppr超能文献

分子序列准确性与蛋白质编码区分析

Molecular sequence accuracy and the analysis of protein coding regions.

作者信息

States D J, Botstein D

机构信息

National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894.

出版信息

Proc Natl Acad Sci U S A. 1991 Jul 1;88(13):5518-22. doi: 10.1073/pnas.88.13.5518.

Abstract

Molecular sequences, like all experimental data, have finite error rates. The impact of errors on the information content of molecular sequence data is dependent on the analytic paradigm used to interpret the data. We studied the impact of nucleic acid sequence errors on the ability to align predicted amino acid sequences with the sequences of related proteins. We found that with a simultaneous translation and alignment algorithm, identification of sequence homologies is resilient to the introduction of random errors. Proteins with greater than 30% sequence identity can be reliably recognized even in the presence of 1% frameshifting (insertion or deletion) error rates and 5% base substitution rates. Incorporation of prior knowledge about the location and characteristics of errors improves tolerance to error of amino acid sequence alignments. Similarly, inclusion of prior knowledge of biased codon utilization by yeast (Saccharomyces cerevisiae) allows reliable detection of correct reading frames in yeast sequences even in the presence of 5% substitution and 1% frameshift errors.

摘要

与所有实验数据一样,分子序列具有有限的错误率。错误对分子序列数据信息内容的影响取决于用于解释数据的分析范式。我们研究了核酸序列错误对将预测的氨基酸序列与相关蛋白质序列进行比对能力的影响。我们发现,使用同步翻译和比对算法时,序列同源性的识别对随机错误的引入具有弹性。即使存在1%的移码(插入或缺失)错误率和5%的碱基替换率,序列同一性大于30%的蛋白质也能被可靠识别。纳入有关错误位置和特征的先验知识可提高氨基酸序列比对的错误耐受性。同样,纳入酵母(酿酒酵母)密码子使用偏好的先验知识,即使存在5%的替换和1%的移码错误,也能可靠检测酵母序列中的正确阅读框。

相似文献

1
Molecular sequence accuracy and the analysis of protein coding regions.分子序列准确性与蛋白质编码区分析
Proc Natl Acad Sci U S A. 1991 Jul 1;88(13):5518-22. doi: 10.1073/pnas.88.13.5518.
7
A tool for multiple sequence alignment.一种用于多序列比对的工具。
Proc Natl Acad Sci U S A. 1989 Jun;86(12):4412-5. doi: 10.1073/pnas.86.12.4412.

引用本文的文献

8
Having a BLAST with bioinformatics (and avoiding BLASTphemy).享受生物信息学带来的乐趣(并避免亵渎生物信息学)。
Genome Biol. 2001;2(10):REVIEWS2002. doi: 10.1186/gb-2001-2-10-reviews2002. Epub 2001 Sep 27.

本文引用的文献

1
Comparative biosequence metrics.比较生物序列度量
J Mol Evol. 1981;18(1):38-46. doi: 10.1007/BF01733210.
2
Identification of common molecular subsequences.常见分子子序列的鉴定
J Mol Biol. 1981 Mar 25;147(1):195-7. doi: 10.1016/0022-2836(81)90087-5.
4
Recognition of protein coding regions in DNA sequences.DNA序列中蛋白质编码区域的识别。
Nucleic Acids Res. 1982 Sep 11;10(17):5303-18. doi: 10.1093/nar/10.17.5303.
5
Establishing homologies in protein sequences.确定蛋白质序列中的同源性。
Methods Enzymol. 1983;91:524-45. doi: 10.1016/s0076-6879(83)91049-2.
9
Primary structure of human neutrophil elastase.人中性粒细胞弹性蛋白酶的一级结构。
Proc Natl Acad Sci U S A. 1987 Apr;84(8):2228-32. doi: 10.1073/pnas.84.8.2228.
10
Multiplex DNA sequencing.多重DNA测序
Science. 1988 Apr 8;240(4849):185-8. doi: 10.1126/science.3353714.

文献AI研究员

20分钟写一篇综述,助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型,支持多种主流文档格式。

立即体验