• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

PGPointNovo:一种用于并行肽段测序的高效神经网络工具。

PGPointNovo: an efficient neural network-based tool for parallel peptide sequencing.

作者信息

Xu Xiaofang, Yang Chunde, He Qiang, Shu Kunxian, Xinpu Yuan, Chen Zhiguang, Zhu Yunping, Chen Tao

机构信息

The School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing 400065, China.

School of Software and Electrical Engineering, Swinburne University of Technology, Melbourne, Victoria 3122, Australia.

出版信息

Bioinform Adv. 2023 Apr 25;3(1):vbad057. doi: 10.1093/bioadv/vbad057. eCollection 2023.

DOI:10.1093/bioadv/vbad057
PMID:37128577
原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC10148685/
Abstract

SUMMARY

peptide sequencing for tandem mass spectrometry data is not only a key technology for novel peptide identification, but also a precedent task for many downstream tasks, such as vaccine and antibody studies. In recent years, neural network models for peptide sequencing have manifested a remarkable ability to accommodate various data sources and outperformed conventional peptide identification tools. However, the excellent model is computationally expensive, taking up to 1 week to process about 400 000 spectrums. This article presents PGPointNovo, a novel neural network-based tool for parallel peptide sequencing. PGPointNovo uses data parallelization technology to accelerate training and inference and optimizes the training obstacles caused by large batch sizes. The results of extensive experiments conducted on multiple datasets of different sizes demonstrate that compared with PointNovo the excellent neural network-based peptide sequencing tool, PGPointNovo, accelerates peptide sequencing by up to 7.35× without precision or recall compromises.

AVAILABILITY AND IMPLEMENTATION

The source code and the parameter settings are available at https://github.com/shallFun4Learning/PGPointNovo.

SUPPLEMENTARY INFORMATION

Supplementary data are available at online.

摘要

摘要

串联质谱数据的肽段测序不仅是鉴定新型肽段的关键技术,也是许多下游任务(如疫苗和抗体研究)的前置任务。近年来,用于肽段测序的神经网络模型已展现出整合各种数据源的卓越能力,且性能优于传统的肽段鉴定工具。然而,性能优异的模型计算成本高昂,处理约400000个谱图耗时长达1周。本文介绍了PGPointNovo,一种基于神经网络的新型并行肽段测序工具。PGPointNovo采用数据并行化技术加速训练和推理,并优化了由大批量数据导致的训练障碍。在多个不同规模数据集上进行的大量实验结果表明,与基于神经网络的优秀肽段测序工具PointNovo相比,PGPointNovo在不损失精度或召回率的情况下,将肽段测序速度提高了7.35倍。

可用性与实现方式

源代码和参数设置可在https://github.com/shallFun4Learning/PGPointNovo获取。

补充信息

补充数据可在线获取。

相似文献

1
PGPointNovo: an efficient neural network-based tool for parallel peptide sequencing.PGPointNovo:一种用于并行肽段测序的高效神经网络工具。
Bioinform Adv. 2023 Apr 25;3(1):vbad057. doi: 10.1093/bioadv/vbad057. eCollection 2023.
2
MRUniNovo: an efficient tool for de novo peptide sequencing utilizing the hadoop distributed computing framework.MRUniNovo:一种利用Hadoop分布式计算框架进行从头肽测序的高效工具。
Bioinformatics. 2017 Mar 15;33(6):944-946. doi: 10.1093/bioinformatics/btw721.
3
SW-Tandem: a highly efficient tool for large-scale peptide identification with parallel spectrum dot product on Sunway TaihuLight.SW-Tandem:在神威·太湖之光上通过并行谱点积进行大规模肽段鉴定的高效工具。
Bioinformatics. 2019 Oct 1;35(19):3861-3863. doi: 10.1093/bioinformatics/btz147.
4
SWPepNovo: An Efficient De Novo Peptide Sequencing Tool for Large-scale MS/MS Spectra Analysis.SWPepNovo:一种用于大规模 MS/MS 谱分析的高效从头肽测序工具。
Int J Biol Sci. 2019 Jul 3;15(9):1787-1801. doi: 10.7150/ijbs.32142. eCollection 2019.
5
DNMSO; an ontology for representing de novo sequencing results from Tandem-MS data.DNMSO:一种用于表示串联质谱数据从头测序结果的本体。
PeerJ. 2020 Oct 21;8:e10216. doi: 10.7717/peerj.10216. eCollection 2020.
6
De novo peptide sequencing by deep learning.通过深度学习进行从头肽测序。
Proc Natl Acad Sci U S A. 2017 Aug 1;114(31):8247-8252. doi: 10.1073/pnas.1705691114. Epub 2017 Jul 18.
7
Folic acid supplementation and malaria susceptibility and severity among people taking antifolate antimalarial drugs in endemic areas.在流行地区,服用抗叶酸抗疟药物的人群中,叶酸补充剂与疟疾易感性和严重程度的关系。
Cochrane Database Syst Rev. 2022 Feb 1;2(2022):CD014217. doi: 10.1002/14651858.CD014217.
8
PDV: an integrative proteomics data viewer.PDV:一种综合蛋白质组学数据查看器。
Bioinformatics. 2019 Apr 1;35(7):1249-1251. doi: 10.1093/bioinformatics/bty770.
9
pNovo 3: precise de novo peptide sequencing using a learning-to-rank framework.pNovo 3:基于排序学习的从头肽序列精确分析。
Bioinformatics. 2019 Jul 15;35(14):i183-i190. doi: 10.1093/bioinformatics/btz366.
10
MCtandem: an efficient tool for large-scale peptide identification on many integrated core (MIC) architecture.MCtandem:一种在许多集成核心 (MIC) 架构上进行大规模肽鉴定的高效工具。
BMC Bioinformatics. 2019 Jul 17;20(1):397. doi: 10.1186/s12859-019-2980-5.

引用本文的文献

1
A learned score function improves the power of mass spectrometry database search.一个有学问的评分函数提高了质谱数据库搜索的能力。
Bioinformatics. 2024 Jun 28;40(Suppl 1):i410-i417. doi: 10.1093/bioinformatics/btae218.
2
Emerging potential of immunopeptidomics by mass spectrometry in cancer immunotherapy.免疫肽组学通过质谱技术在癌症免疫治疗中的新兴潜力。
Cancer Sci. 2024 Apr;115(4):1048-1059. doi: 10.1111/cas.16118. Epub 2024 Feb 21.

本文引用的文献

1
iProX in 2021: connecting proteomics data sharing with big data.iProX 在 2021 年:将蛋白质组学数据共享与大数据连接起来。
Nucleic Acids Res. 2022 Jan 7;50(D1):D1522-D1527. doi: 10.1093/nar/gkab1081.
2
sequencing of proteins by mass spectrometry.质谱法对蛋白质进行测序。
Expert Rev Proteomics. 2020 Jul-Aug;17(7-8):595-607. doi: 10.1080/14789450.2020.1831387. Epub 2020 Oct 21.
3
De novo peptide sequencing by deep learning.通过深度学习进行从头肽测序。
Proc Natl Acad Sci U S A. 2017 Aug 1;114(31):8247-8252. doi: 10.1073/pnas.1705691114. Epub 2017 Jul 18.
4
MRUniNovo: an efficient tool for de novo peptide sequencing utilizing the hadoop distributed computing framework.MRUniNovo:一种利用Hadoop分布式计算框架进行从头肽测序的高效工具。
Bioinformatics. 2017 Mar 15;33(6):944-946. doi: 10.1093/bioinformatics/btw721.