美国国立生物技术信息中心参考序列（RefSeq）：一个经过整理的基因组、转录本和蛋白质的非冗余序列数据库。

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

作者信息

Pruitt Kim D, Tatusova Tatiana, Maglott Donna R

机构信息

National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Rm 6An.12J, 45 Center Drive, Bethesda, MD 20892-6510, USA.

出版信息

Nucleic Acids Res. 2007 Jan;35(Database issue):D61-5. doi: 10.1093/nar/gkl842. Epub 2006 Nov 27.

DOI:10.1093/nar/gkl842

PMID:17130148

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC1716718/

Abstract

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

摘要

美国国立医学图书馆国家生物技术信息中心（NCBI）的参考序列（RefSeq）数据库（http://www.ncbi.nlm.nih.gov/RefSeq/）是一个经过整理的非冗余序列集合，涵盖基因组、转录本和蛋白质。该数据库包含3774种生物，涵盖原核生物、真核生物和病毒，拥有2879860条蛋白质记录（RefSeq第19版）。RefSeq记录整合了来自多个来源的信息，当这些来源有额外数据时，因此代表了序列及其特征的当前描述。注释包括编码区、保守结构域、tRNA、序列标签位点（STS）、变异、参考文献、基因和蛋白质产物名称以及数据库交叉引用。序列经过审核，并使用合作及科学界其他输入、预测、从GenBank传播以及NCBI工作人员整理的组合方法添加特征。所有RefSeq记录的格式都经过验证，并且正在应用越来越多的测试来评估序列和注释的质量，特别是在完整基因组序列的背景下。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/dedc/1781170/a382f8994184/gkl842f1.jpg

相似文献

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

Nucleic Acids Res. 2007 Jan;35(Database issue):D61-5. doi: 10.1093/nar/gkl842. Epub 2006 Nov 27.

NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4. doi: 10.1093/nar/gki025.

NCBI Reference Sequences (RefSeq): current status, new features and genome annotation policy.

Nucleic Acids Res. 2012 Jan;40(Database issue):D130-5. doi: 10.1093/nar/gkr1079. Epub 2011 Nov 24.

NCBI Reference Sequences: current status, policy and new initiatives.

Nucleic Acids Res. 2009 Jan;37(Database issue):D32-6. doi: 10.1093/nar/gkn721. Epub 2008 Oct 16.

Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation.

Nucleic Acids Res. 2016 Jan 4;44(D1):D733-45. doi: 10.1093/nar/gkv1189. Epub 2015 Nov 8.

Comparison of RefSeq protein-coding regions in human and vertebrate genomes.

BMC Genomics. 2013 Sep 25;14:654. doi: 10.1186/1471-2164-14-654.

Gene: a gene-centered information resource at NCBI.

Nucleic Acids Res. 2015 Jan;43(Database issue):D36-42. doi: 10.1093/nar/gku1055. Epub 2014 Oct 29.

RefSeq: an update on mammalian reference sequences.

Nucleic Acids Res. 2014 Jan;42(Database issue):D756-63. doi: 10.1093/nar/gkt1114. Epub 2013 Nov 19.

Entrez Gene: gene-centered information at NCBI.

Nucleic Acids Res. 2011 Jan;39(Database issue):D52-7. doi: 10.1093/nar/gkq1237. Epub 2010 Nov 28.

Entrez Gene: gene-centered information at NCBI.

Nucleic Acids Res. 2007 Jan;35(Database issue):D26-31. doi: 10.1093/nar/gkl993. Epub 2006 Dec 5.

引用本文的文献

Non-coding RNA profiling in BRAF-mutant cutaneous melanoma before and after Spry1 depletion.

Sci Data. 2025 Sep 2;12(1):1538. doi: 10.1038/s41597-025-05807-x.

Chromosome-level genome assembly of the caddisfly Stenopsyche angustata (Insecta: Trichoptera).

Sci Data. 2025 Sep 1;12(1):1523. doi: 10.1038/s41597-025-05602-8.

Leveraging the enrichment analysis from a genome-wide association study against epilepsy-focusing on the role of tryptophan catabolites pathway in patients with drug-resistant epilepsy.

Front Nutr. 2025 Aug 6;12:1539145. doi: 10.3389/fnut.2025.1539145. eCollection 2025.

Transcriptome sequencing reveals the evolutionary histories and gene expression evolution in two related Pagurus species.

PLoS One. 2025 Aug 20;20(8):e0330170. doi: 10.1371/journal.pone.0330170. eCollection 2025.

Analysis wheat wild relatives Thinopyrum intermedium and Roegneria kamoji genomes reveal different polyploid evolution paths.

Nat Commun. 2025 Aug 18;16(1):7693. doi: 10.1038/s41467-025-63007-y.

LDAK-KVIK performs fast and powerful mixed-model association analysis of quantitative and binary phenotypes.

Nat Genet. 2025 Aug 11. doi: 10.1038/s41588-025-02286-z.

Metagenomic analysis reveals methanogenic and other archaeal genes in the digestive tract of invasive Japanese beetle larvae and associated soil.

Front Microbiol. 2025 Jul 25;16:1609893. doi: 10.3389/fmicb.2025.1609893. eCollection 2025.

Bag-of-words is competitive with sum-of-embeddings language-inspired representations on protein inference.

PLoS One. 2025 Aug 6;20(8):e0325531. doi: 10.1371/journal.pone.0325531. eCollection 2025.

Ancestral Sequence Reconstruction of the Ethylene-Forming Enzyme.

Biochemistry. 2025 Aug 5;64(15):3432-3445. doi: 10.1021/acs.biochem.5c00334. Epub 2025 Jul 25.

Harnessing Nanobodies for Precision Targeting of Proteoforms: Opportunities and Challenges in Therapeutics and Diagnostics.

ACS Chem Biol. 2025 Aug 15;20(8):1817-1827. doi: 10.1021/acschembio.5c00329. Epub 2025 Jul 24.

本文引用的文献

Entrez Gene: gene-centered information at NCBI.

Nucleic Acids Res. 2011 Jan;39(Database issue):D52-7. doi: 10.1093/nar/gkq1237. Epub 2010 Nov 28.

Database resources of the National Center for Biotechnology Information.

Nucleic Acids Res. 2007 Jan;35(Database issue):D5-12. doi: 10.1093/nar/gkl1031. Epub 2006 Dec 14.

The Mouse Genome Database (MGD): updates and enhancements.

Nucleic Acids Res. 2006 Jan 1;34(Database issue):D562-7. doi: 10.1093/nar/gkj085.

WormBase: better software, richer content.

Nucleic Acids Res. 2006 Jan 1;34(Database issue):D475-8. doi: 10.1093/nar/gkj061.

FlyBase: genes and gene models.

Nucleic Acids Res. 2005 Jan 1;33(Database issue):D390-5. doi: 10.1093/nar/gki046.

Regulation of gene expression by stop codon recoding: selenocysteine.

Gene. 2003 Jul 17;312:17-25. doi: 10.1016/s0378-1119(03)00588-2.

Generation of protein isoform diversity by alternative initiation of translation at non-AUG codons.

Biol Cell. 2003 May-Jun;95(3-4):169-78. doi: 10.1016/s0248-4900(03)00033-9.

The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community.

Nucleic Acids Res. 2003 Jan 1;31(1):224-8. doi: 10.1093/nar/gkg076.

GenBank.

Nucleic Acids Res. 2000 Jan 1;28(1):15-8. doi: 10.1093/nar/28.1.15.

Complete genomes in WWW Entrez: data representation and analysis.

Bioinformatics. 1999 Jul-Aug;15(7-8):536-43. doi: 10.1093/bioinformatics/15.7.536.

文献AI研究员

20分钟写一篇综述，助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型，支持多种主流文档格式。

立即体验

美国国立生物技术信息中心参考序列（RefSeq）：一个经过整理的基因组、转录本和蛋白质的非冗余序列数据库。

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

作者信息

机构信息

出版信息

相似文献

引用本文的文献

本文引用的文献

文献AI研究员

用中文搜PubMed

文档翻译

Suppr 超能文献

美国国立生物技术信息中心参考序列（RefSeq）：一个经过整理的基因组、转录本和蛋白质的非冗余序列数据库。

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

作者信息

机构信息

出版信息

相似文献

引用本文的文献

本文引用的文献