Suppr超能文献

利用信息论和松弛方法对无间隙DNA序列进行快速多重比对。

Fast multiple alignment of ungapped DNA sequences using information theory and a relaxation method.

作者信息

Schneider Thomas D, Mastronarde David N

机构信息

National Cancer Institute, Frederick Cancer Research and Development Center, Laboratory of Mathematical Biology, P. O. Box B, Frederick, MD 21702-1201.

出版信息

Discrete Appl Math. 1996 Dec 1;71(1-3):259-268. doi: 10.1016/S0166-218X(96)00068-6.

Abstract

An information theory based multiple alignment ("Malign") method was used to align the DNA binding sequences of the OxyR and Fis proteins, whose sequence conservation is so spread out that it is difficult to identify the sites. In the algorithm described here, the information content of the sequences is used as a unique global criterion for the quality of the alignment. The algorithm uses look-up tables to avoid recalculating computationally expensive functions such as the logarithm. Because there are no arbitrary constants and because the results are reported in absolute units (bits), the best alignment can be chosen without ambiguity. Starting from randomly selected alignments, a hill-climbing algorithm can track through the immense space of s(n) combinations where s is the number of sequences and n is the number of positions possible for each sequence. Instead of producing a single alignment, the algorithm is fast enough that one can afford to use many start points and to classify the solutions. Good convergence is indicated by the presence of a single well-populated solution class having higher information content than other classes. The existence of several distinct classes for the Fis protein indicates that those binding sites have self-similar features.

摘要

一种基于信息论的多重比对(“Malign”)方法被用于比对OxyR和Fis蛋白的DNA结合序列,这些序列的保守性分布得非常分散,以至于难以识别位点。在此描述的算法中,序列的信息含量被用作比对质量的唯一全局标准。该算法使用查找表来避免重新计算计算成本高昂的函数,如对数函数。由于没有任意常数,并且结果以绝对单位(比特)报告,因此可以明确无误地选择最佳比对。从随机选择的比对开始,爬山算法可以在s(n)组合的巨大空间中进行跟踪,其中s是序列的数量,n是每个序列可能的位置数量。该算法不是产生单个比对,而是速度足够快,以至于可以使用许多起始点并对解决方案进行分类。单个信息含量高于其他类别的密集填充的解决方案类别的存在表明收敛良好。Fis蛋白存在几个不同的类别,这表明那些结合位点具有自相似特征。

相似文献

3
Using CLUSTAL for multiple sequence alignments.使用CLUSTAL进行多序列比对。
Methods Enzymol. 1996;266:383-402. doi: 10.1016/s0076-6879(96)66024-8.
7
Fast model-based protein homology detection without alignment.基于快速模型的无需比对的蛋白质同源性检测。
Bioinformatics. 2007 Jul 15;23(14):1728-36. doi: 10.1093/bioinformatics/btm247. Epub 2007 May 8.

引用本文的文献

1
Analysis of plant metabolomics data using identification-free approaches.使用无鉴定方法分析植物代谢组学数据。
Appl Plant Sci. 2025 Mar 1;13(4):e70001. doi: 10.1002/aps3.70001. eCollection 2025 Jul-Aug.
4
Trends in information theory-based chemical structure codification.基于信息论的化学结构编码趋势。
Mol Divers. 2014 Aug;18(3):673-86. doi: 10.1007/s11030-014-9517-7. Epub 2014 Apr 5.

本文引用的文献

4
A multiple sequence comparison method.一种多序列比对方法。
Bull Math Biol. 1993 Mar;55(2):465-86. doi: 10.1007/BF02460892.
6
Delila system tools.德利拉系统工具。
Nucleic Acids Res. 1984 Jan 11;12(1 Pt 1):129-40. doi: 10.1093/nar/12.1part1.129.
8

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验