• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

SHI7是一种用于多用途短读长DNA质量控制的自学习流程。

SHI7 Is a Self-Learning Pipeline for Multipurpose Short-Read DNA Quality Control.

作者信息

Al-Ghalith Gabriel A, Hillmann Benjamin, Ang Kaiwei, Shields-Cutler Robin, Knights Dan

机构信息

Bioinformatics and Computational Biology, University of Minnesota-Twin Cities, Minneapolis, Minnesota, USA.

Computer Science, University of Minnesota-Twin Cities, Minneapolis, Minnesota, USA.

出版信息

mSystems. 2018 Apr 24;3(3). doi: 10.1128/mSystems.00202-17. eCollection 2018 May-Jun.

DOI:10.1128/mSystems.00202-17
PMID:29719872
原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC5915699/
Abstract

Next-generation sequencing technology is of great importance for many biological disciplines; however, due to technical and biological limitations, the short DNA sequences produced by modern sequencers require numerous quality control (QC) measures to reduce errors, remove technical contaminants, or merge paired-end reads together into longer or higher-quality contigs. Many tools for each step exist, but choosing the appropriate methods and usage parameters can be challenging because the parameterization of each step depends on the particularities of the sequencing technology used, the type of samples being analyzed, and the stochasticity of the instrumentation and sample preparation. Furthermore, end users may not know all of the relevant information about how their data were generated, such as the expected overlap for paired-end sequences or type of adaptors used to make informed choices. This increasing complexity and nuance demand a pipeline that combines existing steps together in a user-friendly way and, when possible, learns reasonable quality parameters from the data automatically. We propose a user-friendly quality control pipeline called SHI7 (canonically pronounced "shizen"), which aims to simplify quality control of short-read data for the end user by predicting presence and/or type of common sequencing adaptors, what quality scores to trim, whether the data set is shotgun or amplicon sequencing, whether reads are paired end or single end, and whether pairs are stitchable, including the expected amount of pair overlap. We hope that SHI7 will make it easier for all researchers, expert and novice alike, to follow reasonable practices for short-read data quality control. Quality control of high-throughput DNA sequencing data is an important but sometimes laborious task requiring background knowledge of the sequencing protocol used (such as adaptor type, sequencing technology, insert size/stitchability, paired-endedness, etc.). Quality control protocols typically require applying this background knowledge to selecting and executing numerous quality control steps with the appropriate parameters, which is especially difficult when working with public data or data from collaborators who use different protocols. We have created a streamlined quality control pipeline intended to substantially simplify the process of DNA quality control from raw machine output files to actionable sequence data. In contrast to other methods, our proposed pipeline is easy to install and use and attempts to learn the necessary parameters from the data automatically with a single command.

摘要

下一代测序技术对许多生物学学科都非常重要;然而,由于技术和生物学上的限制,现代测序仪产生的短DNA序列需要众多质量控制(QC)措施来减少错误、去除技术污染物,或将双端读数合并成更长或质量更高的重叠群。针对每个步骤都有许多工具,但选择合适的方法和使用参数可能具有挑战性,因为每个步骤的参数设置取决于所使用的测序技术的特殊性、被分析样本的类型以及仪器和样本制备的随机性。此外,终端用户可能并不了解有关其数据如何生成的所有相关信息,例如双端序列的预期重叠或用于做出明智选择的接头类型。这种日益增加的复杂性和细微差别需要一个以用户友好的方式将现有步骤组合在一起的流程,并且在可能的情况下,能从数据中自动学习合理的质量参数。我们提出了一个名为SHI7(标准发音为“shizen”)的用户友好型质量控制流程,其目的是通过预测常见测序接头的存在和/或类型、要修剪的质量分数、数据集是鸟枪法测序还是扩增子测序、读数是双端还是单端以及双端是否可拼接(包括预期的双端重叠量),来为终端用户简化短读数据的质量控制。我们希望SHI7能让所有研究人员,无论是专家还是新手,都更容易遵循短读数据质量控制的合理做法。高通量DNA测序数据的质量控制是一项重要但有时很费力的任务,需要对所使用的测序方案有背景知识(如接头类型、测序技术、插入片段大小/可拼接性、双端性等)。质量控制方案通常需要应用这些背景知识来选择并执行众多具有适当参数的质量控制步骤,在处理公共数据或来自使用不同方案的合作者的数据时尤其困难。我们创建了一个简化的质量控制流程,旨在从原始机器输出文件到可操作的序列数据,大幅简化DNA质量控制过程。与其他方法不同,我们提出的流程易于安装和使用,并尝试通过单个命令从数据中自动学习必要的参数。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1a19/5915699/10544e466a1f/sys0031822260003.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1a19/5915699/d3ca18130723/sys0031822260001.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1a19/5915699/377fb89865c8/sys0031822260002.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1a19/5915699/10544e466a1f/sys0031822260003.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1a19/5915699/d3ca18130723/sys0031822260001.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1a19/5915699/377fb89865c8/sys0031822260002.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/1a19/5915699/10544e466a1f/sys0031822260003.jpg

相似文献

1
SHI7 Is a Self-Learning Pipeline for Multipurpose Short-Read DNA Quality Control.SHI7是一种用于多用途短读长DNA质量控制的自学习流程。
mSystems. 2018 Apr 24;3(3). doi: 10.1128/mSystems.00202-17. eCollection 2018 May-Jun.
2
CDSnake: Snakemake pipeline for retrieval of annotated OTUs from paired-end reads using CD-HIT utilities.CDSnake:使用 CD-HIT 工具从配对末端读取中检索带注释的 OTU 的 Snakemake 管道。
BMC Bioinformatics. 2020 Jul 24;21(Suppl 12):303. doi: 10.1186/s12859-020-03591-6.
3
Benefits of merging paired-end reads before pre-processing environmental metagenomics data.在预处理环境宏基因组数据之前合并配对末端reads 的好处。
Mar Genomics. 2022 Feb;61:100914. doi: 10.1016/j.margen.2021.100914. Epub 2021 Dec 2.
4
ClinQC: a tool for quality control and cleaning of Sanger and NGS data in clinical research.ClinQC:临床研究中用于Sanger测序和二代测序(NGS)数据质量控制与清理的工具
BMC Bioinformatics. 2016 Feb 2;17:56. doi: 10.1186/s12859-016-0915-y.
5
Benchmarking software tools for trimming adapters and merging next-generation sequencing data for ancient DNA.用于修剪接头和合并古代DNA下一代测序数据的基准测试软件工具。
Front Bioinform. 2023 Dec 7;3:1260486. doi: 10.3389/fbinf.2023.1260486. eCollection 2023.
6
ASAP 2: a pipeline and web server to analyze marker gene amplicon sequencing data automatically and consistently.ASAP 2:一个用于自动和一致地分析标记基因扩增子测序数据的流水线和网络服务器。
BMC Bioinformatics. 2022 Jan 6;23(1):27. doi: 10.1186/s12859-021-04555-0.
7
Unlocking short read sequencing for metagenomics.解锁宏基因组学的短读测序。
PLoS One. 2010 Jul 28;5(7):e11840. doi: 10.1371/journal.pone.0011840.
8
Contig-Layout-Authenticator (CLA): A Combinatorial Approach to Ordering and Scaffolding of Bacterial Contigs for Comparative Genomics and Molecular Epidemiology.重叠群布局验证器(CLA):一种用于比较基因组学和分子流行病学的细菌重叠群排序与搭建支架的组合方法。
PLoS One. 2016 Jun 1;11(6):e0155459. doi: 10.1371/journal.pone.0155459. eCollection 2016.
9
Graph mining for next generation sequencing: leveraging the assembly graph for biological insights.用于下一代测序的图挖掘:利用组装图获取生物学见解。
BMC Genomics. 2016 May 6;17:340. doi: 10.1186/s12864-016-2678-2.
10
Computational Pipeline for the Detection of Plant RNA Viruses Using High-Throughput Sequencing.基于高通量测序的植物 RNA 病毒检测计算流程。
Methods Mol Biol. 2024;2724:1-20. doi: 10.1007/978-1-0716-3485-1_1.

引用本文的文献

1
Extensive novel diversity and phenotypic associations in the dromedary camel microbiome are revealed through deep metagenomics and machine learning.通过深度宏基因组学和机器学习揭示了单峰骆驼微生物组中广泛的新多样性和表型关联。
PLoS One. 2025 Jul 17;20(7):e0328194. doi: 10.1371/journal.pone.0328194. eCollection 2025.
2
Gut mycobiome maturation and its determinants during early childhood: a comparison of ITS2 amplicon and shotgun metagenomic sequencing approaches.幼儿期肠道真菌群落成熟及其决定因素:ITS2扩增子测序与鸟枪法宏基因组测序方法的比较
Front Microbiol. 2025 May 21;16:1539750. doi: 10.3389/fmicb.2025.1539750. eCollection 2025.
3

本文引用的文献

1
Microbiome Helper: a Custom and Streamlined Workflow for Microbiome Research.微生物组助手:一种用于微生物组研究的定制且简化的工作流程。
mSystems. 2017 Jan 3;2(1). doi: 10.1128/mSystems.00127-16. eCollection 2017 Jan-Feb.
2
Captivity humanizes the primate microbiome.圈养使灵长类动物的微生物群更具人类特征。
Proc Natl Acad Sci U S A. 2016 Sep 13;113(37):10376-81. doi: 10.1073/pnas.1521835113. Epub 2016 Aug 29.
3
NINJA-OPS: Fast Accurate Marker Gene Alignment Using Concatenated Ribosomes.忍者行动:使用串联核糖体进行快速准确的标记基因比对
A Randomized Pilot Study of Time-Restricted Eating Shows Minimal Microbiome Changes.
一项关于限时进食的随机试点研究表明微生物组变化极小。
Nutrients. 2025 Jan 4;17(1):185. doi: 10.3390/nu17010185.
4
Maternal oral probiotic use is associated with decreased breastmilk inflammatory markers, infant fecal microbiome variation, and altered recognition memory responses in infants-a pilot observational study.孕妇口服益生菌与母乳炎症标志物减少、婴儿粪便微生物群变化以及婴儿识别记忆反应改变相关——一项初步观察性研究。
Front Nutr. 2024 Sep 25;11:1456111. doi: 10.3389/fnut.2024.1456111. eCollection 2024.
5
Human milk variation is shaped by maternal genetics and impacts the infant gut microbiome.人乳的变化受母体遗传影响,并影响婴儿肠道微生物组。
Cell Genom. 2024 Oct 9;4(10):100638. doi: 10.1016/j.xgen.2024.100638. Epub 2024 Sep 11.
6
Human cytomegalovirus in breast milk is associated with milk composition and the infant gut microbiome and growth.人巨细胞病毒在母乳中与乳汁成分以及婴儿肠道微生物组和生长有关。
Nat Commun. 2024 Jul 23;15(1):6216. doi: 10.1038/s41467-024-50282-4.
7
A Randomized Controlled Trial to Evaluate the Impact of a Novel Probiotic and Nutraceutical Supplement on Pruritic Dermatitis and the Gut Microbiota in Privately Owned Dogs.一项随机对照试验,旨在评估一种新型益生菌和营养补充剂对宠物狗瘙痒性皮炎和肠道微生物群的影响。
Animals (Basel). 2024 Jan 30;14(3):453. doi: 10.3390/ani14030453.
8
Variovorax sp. strain P1R9 applied individually or as part of bacterial consortia enhances wheat germination under salt stress conditions.聚生镰孢菌 P1R9 单独或作为细菌混合体的一部分应用可增强小麦在盐胁迫条件下的萌发。
Sci Rep. 2024 Jan 24;14(1):2070. doi: 10.1038/s41598-024-52535-0.
9
Endophytic bacterial communities in ungerminated and germinated seeds of commercial vegetables.商业蔬菜未发芽和发芽种子中的内生细菌群落。
Sci Rep. 2023 Nov 14;13(1):19829. doi: 10.1038/s41598-023-47099-4.
10
Bacterial communities associated with wood rot fungi that use distinct decomposition mechanisms.与采用不同分解机制的木腐真菌相关的细菌群落。
ISME Commun. 2022 Mar 30;2(1):26. doi: 10.1038/s43705-022-00108-5.
PLoS Comput Biol. 2016 Jan 28;12(1):e1004658. doi: 10.1371/journal.pcbi.1004658. eCollection 2016 Jan.
4
Trimmomatic: a flexible trimmer for Illumina sequence data.Trimmomatic:一款适用于 Illumina 测序数据的灵活修剪工具。
Bioinformatics. 2014 Aug 1;30(15):2114-20. doi: 10.1093/bioinformatics/btu170. Epub 2014 Apr 1.
5
Using QIIME to analyze 16S rRNA gene sequences from microbial communities.使用QIIME分析来自微生物群落的16S rRNA基因序列。
Curr Protoc Microbiol. 2012 Nov;Chapter 1:Unit 1E.5.. doi: 10.1002/9780471729259.mc01e05s27.
6
Structure, function and diversity of the healthy human microbiome.健康人体微生物组的结构、功能与多样性。
Nature. 2012 Jun 13;486(7402):207-14. doi: 10.1038/nature11234.
7
Fast gapped-read alignment with Bowtie 2.快速缺口读对准与 Bowtie 2。
Nat Methods. 2012 Mar 4;9(4):357-9. doi: 10.1038/nmeth.1923.
8
Experimental and analytical tools for studying the human microbiome.研究人类微生物组的实验和分析工具。
Nat Rev Genet. 2011 Dec 16;13(1):47-58. doi: 10.1038/nrg3129.
9
FLASH: fast length adjustment of short reads to improve genome assemblies.FLASH:快速调整短读长以提高基因组组装质量。
Bioinformatics. 2011 Nov 1;27(21):2957-63. doi: 10.1093/bioinformatics/btr507. Epub 2011 Sep 7.
10
Search and clustering orders of magnitude faster than BLAST.比 BLAST 快几个数量级的搜索和聚类。
Bioinformatics. 2010 Oct 1;26(19):2460-1. doi: 10.1093/bioinformatics/btq461. Epub 2010 Aug 12.