• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

用于先进蛋白质工程的统一序列结构编码——一种多模态扩散变换器。

Unifying sequence-structure coding for advanced protein engineering a multimodal diffusion transformer.

作者信息

Lin Xiaohan, Chen Zhenyu, Li Yanheng, Ma Zicheng, Fan Chuanliu, Cao Ziqiang, Feng Shihao, Zhang Jun, Gao Yi Qin

机构信息

Beijing National Laboratory for Molecular Sciences, College of Chemistry and Molecular Engineering, Peking University Beijing 100871 China

Changping Laboratory Beijing 102200 China

出版信息

Chem Sci. 2025 May 15. doi: 10.1039/d5sc02055g.

DOI:10.1039/d5sc02055g
PMID:40417294
原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC12096517/
Abstract

Modern protein engineering demands integrated sequence-structure representations to tackle key challenges in designing, modifying, and evolving proteins for specific functions. While sequence-based methods are promising for generating novel proteins, incorporating structure-oriented information improves the success rate and helps target corresponding functions. Therefore, rather than relying solely on sequence or structure-based approaches, a consensus strategy is essential. Here, we introduce ProTokens, machine-learned "amino acids" derived from structural databases self-supervised learning, providing a compact yet information-rich representation that bridges sequence and structure modalities. Instead of treating sequences and structures separately, we build PT-DiT, a multimodal diffusion transformer-based model that integrates both into a unified representation, enabling protein engineering in a joint sequence-structure space, streamlining the design process and facilitating the efficient encoding of 3D folds, contextual protein design, sampling of metastable states, and directed evolution for diverse objectives. Therefore, as a unified solution for protein engineering, PT-DiT leverages sequence and structure insights to realize functional protein design.

摘要

现代蛋白质工程需要整合的序列-结构表示,以应对在设计、修饰和进化具有特定功能的蛋白质方面的关键挑战。虽然基于序列的方法有望生成新型蛋白质,但纳入面向结构的信息可提高成功率并有助于靶向相应功能。因此,共识策略至关重要,而不是仅仅依赖基于序列或结构的方法。在这里,我们引入了ProTokens,这是一种通过自监督学习从结构数据库中衍生出来的机器学习“氨基酸”,它提供了一种紧凑但信息丰富的表示,架起了序列和结构模态之间的桥梁。我们没有分别处理序列和结构,而是构建了PT-DiT,这是一个基于多模态扩散变换器的模型,它将两者整合到一个统一的表示中,能够在联合序列-结构空间中进行蛋白质工程,简化设计过程,并促进三维折叠的高效编码、上下文蛋白质设计、亚稳态采样以及针对不同目标的定向进化。因此,作为蛋白质工程的统一解决方案,PT-DiT利用序列和结构见解来实现功能性蛋白质设计。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/39a0a8b1726e/d5sc02055g-f5.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/230ed62a2e9f/d5sc02055g-f1.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/e6d1a04eb00d/d5sc02055g-f2.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/19d560ec8f80/d5sc02055g-f3.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/de98e43950bb/d5sc02055g-f4.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/39a0a8b1726e/d5sc02055g-f5.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/230ed62a2e9f/d5sc02055g-f1.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/e6d1a04eb00d/d5sc02055g-f2.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/19d560ec8f80/d5sc02055g-f3.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/de98e43950bb/d5sc02055g-f4.jpg
https://cdn.ncbi.nlm.nih.gov/pmc/blobs/a8d3/12096517/39a0a8b1726e/d5sc02055g-f5.jpg

相似文献

1
Unifying sequence-structure coding for advanced protein engineering a multimodal diffusion transformer.用于先进蛋白质工程的统一序列结构编码——一种多模态扩散变换器。
Chem Sci. 2025 May 15. doi: 10.1039/d5sc02055g.
2
Stakeholders' perceptions and experiences of factors influencing the commissioning, delivery, and uptake of general health checks: a qualitative evidence synthesis.利益相关者对影响一般健康检查的委托、提供和接受因素的看法与体验:一项定性证据综合分析
Cochrane Database Syst Rev. 2025 Mar 20;3(3):CD014796. doi: 10.1002/14651858.CD014796.pub2.
3
Prediction, screening and characterization of novel bioactive tetrapeptide matrikines for skin rejuvenation.预测、筛选和鉴定具有皮肤年轻化功效的新型生物活性四肽基质。
Br J Dermatol. 2024 Jun 20;191(1):92-106. doi: 10.1093/bjd/ljae061.
4
Aural toilet (ear cleaning) for chronic suppurative otitis media.慢性化脓性中耳炎的耳道清理(耳部清洁)
Cochrane Database Syst Rev. 2025 Jun 9;6(6):CD013057. doi: 10.1002/14651858.CD013057.pub3.
5
Assessing the comparative effects of interventions in COPD: a tutorial on network meta-analysis for clinicians.评估慢性阻塞性肺疾病干预措施的比较效果:面向临床医生的网状Meta分析教程
Respir Res. 2024 Dec 21;25(1):438. doi: 10.1186/s12931-024-03056-x.
6
Wood Waste Valorization and Classification Approaches: A systematic review.木材废料的增值与分类方法:一项系统综述
Open Res Eur. 2025 May 6;5:5. doi: 10.12688/openreseurope.18862.1. eCollection 2025.
7
Journals Operating Predatory Practices Are Systematically Eroding the Science Ethos: A Gate and Code Strategy to Minimise Their Operating Space and Restore Research Best Practice.采用掠夺性做法的期刊正在系统性地侵蚀科学精神:一种减少其运营空间并恢复研究最佳实践的把关与编码策略。
Microb Biotechnol. 2025 Jun;18(6):e70180. doi: 10.1111/1751-7915.70180.
8
AI-Driven Antimicrobial Peptide Discovery: Mining and Generation.人工智能驱动的抗菌肽发现:挖掘与生成
Acc Chem Res. 2025 Jun 17;58(12):1831-1846. doi: 10.1021/acs.accounts.0c00594. Epub 2025 Jun 3.
9
exploits host- and bacterial-derived β-alanine for replication inside host macrophages.利用宿主和细菌来源的β-丙氨酸在宿主巨噬细胞内进行复制。
Elife. 2025 Jun 19;13:RP103714. doi: 10.7554/eLife.103714.
10
Surveillance for Violent Deaths - National Violent Death Reporting System, 50 States, the District of Columbia, and Puerto Rico, 2022.暴力死亡监测——2022年全国暴力死亡报告系统,50个州、哥伦比亚特区和波多黎各
MMWR Surveill Summ. 2025 Jun 12;74(5):1-42. doi: 10.15585/mmwr.ss7405a1.

本文引用的文献

1
Accelerated enzyme engineering by machine-learning guided cell-free expression.通过机器学习引导的无细胞表达实现加速酶工程。
Nat Commun. 2025 Jan 20;16(1):865. doi: 10.1038/s41467-024-55399-0.
2
Active learning-assisted directed evolution.主动学习辅助的定向进化
Nat Commun. 2025 Jan 16;16(1):714. doi: 10.1038/s41467-025-55987-8.
3
Simulating 500 million years of evolution with a language model.用语言模型模拟5亿年的进化历程。
Science. 2025 Feb 21;387(6736):850-858. doi: 10.1126/science.ads0018. Epub 2025 Jan 16.
4
Rapid in silico directed evolution by a protein language model with EVOLVEpro.通过带有EVOLVEpro的蛋白质语言模型进行快速计算机辅助定向进化。
Science. 2025 Jan 24;387(6732):eadr6006. doi: 10.1126/science.adr6006.
5
UniProt: the Universal Protein Knowledgebase in 2025.通用蛋白质知识库(UniProt):2025年的情况
Nucleic Acids Res. 2025 Jan 6;53(D1):D609-D617. doi: 10.1093/nar/gkae1010.
6
Multistate and functional protein design using RoseTTAFold sequence space diffusion.使用RoseTTAFold序列空间扩散进行多状态和功能性蛋白质设计。
Nat Biotechnol. 2024 Sep 25. doi: 10.1038/s41587-024-02395-w.
7
AlphaFold predictions of fold-switched conformations are driven by structure memorization.AlphaFold 对构象转换构象的预测是由结构记忆驱动的。
Nat Commun. 2024 Aug 24;15(1):7296. doi: 10.1038/s41467-024-51801-z.
8
Accurate structure prediction of biomolecular interactions with AlphaFold 3.利用 AlphaFold 3 进行生物分子相互作用的精确结构预测。
Nature. 2024 Jun;630(8016):493-500. doi: 10.1038/s41586-024-07487-w. Epub 2024 May 8.
9
Ligand efficacy modulates conformational dynamics of the µ-opioid receptor.配体效能调节μ阿片受体的构象动力学。
Nature. 2024 May;629(8011):474-480. doi: 10.1038/s41586-024-07295-2. Epub 2024 Apr 10.
10
AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences.2024 年的 AlphaFold 蛋白质结构数据库:为超过 2.14 亿个蛋白质序列提供结构覆盖。
Nucleic Acids Res. 2024 Jan 5;52(D1):D368-D375. doi: 10.1093/nar/gkad1011.