• Suppr超能文献
  • 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
定价套餐&价格
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Win 客户端微信小程序
定价
会员套餐积分包API 积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026
  1. 首页
  2. 分享广场
  3. 虚拟细胞的技术原理、应用场景与发展前景综述

虚拟细胞的技术原理、应用场景与发展前景综述

深度研究匿名用户发表于 2026年01月14日 11:543阅读
发起深度研究
发起深度研究

1. 虚拟细胞的概念内涵与核心特征

1.1 虚拟细胞的定义与本质

虚拟细胞(Virtual Cell)本质上是一种多尺度、多模态的生物系统计算模型,旨在通过整合海量的实验数据与先进的计算模拟技术,高保真地复现真实细胞的动态行为和复杂功能 1。它超越了传统的单一分子或通路研究范式,致力于构建一个全面的、可预测的细胞数字孪生体。这种模型能够捕捉从分子相互作用、基因表达、细胞器功能到细胞形态和集体行为等多个层面的生物学过程,并模拟这些过程在不同生理和病理条件下的动态演变 1。

虚拟细胞的核心属性在于其“多尺度”和“多模态”特性。多尺度意味着它能同时模拟和分析从微观的分子级别(如蛋白质相互作用、代谢通路)到宏观的细胞级别(如细胞增殖、分化、迁移)乃至组织级别(如细胞间相互作用、组织重塑)的复杂生物学现象 12。例如,在骨折愈合的研究中,虚拟细胞模型可以整合组织、细胞和分子层面的生物活动,并考虑机械和生物因素的影响 3。多模态则强调虚拟细胞能够整合来自不同实验技术的数据类型,如单细胞测序、空间转录组、蛋白质组学等“组学”数据,以及细胞力学、电生理等功能数据 1。通过这些异构数据的融合,虚拟细胞能够建立起一个动态交互的网络,从而实现对细胞内部信号传导、基因调控、代谢网络以及细胞与微环境之间相互作用的精确模拟和预测 14。

虚拟细胞的最终目标是提供一个可供科学家进行“体外”实验的计算平台,即“虚拟实验室” 25。在这个虚拟环境中,研究人员可以通过调整参数、引入扰动(如药物干预、基因编辑),来观察细胞行为的变化,从而加速生物学发现,指导实验设计,并对细胞功能有更深入的理解 16。这种迭代的“模拟-实验-验证”闭环,是虚拟细胞实现高保真复现和预测能力的关键。

1.2 与传统细胞模型的区别

传统细胞模型,尤其是还原论方法,通常侧重于研究单一的分子、通路或细胞器功能。例如,一个实验可能专注于揭示特定基因的表达调控机制,或者探究某种蛋白质在信号传导中的作用,其核心是 isolating 和 characterizing individual components 7。这类研究方法虽然在解析具体生物学机制方面取得了巨大成功,但往往难以捕捉细胞作为一个整体的复杂性和动态性。它们通常面临以下局限性:

  1. 割裂性与局限性: 传统模型常常将细胞视为一系列独立运行的模块,忽略了各模块之间错综复杂的相互作用和反馈循环。例如,单一分子机制的研究很难解释为何同一个分子在不同细胞类型或微环境下会产生截然不同的生物学效应。
  2. 静态视角: 许多传统模型提供的是细胞在特定时间点或特定条件下的“快照”,难以动态地模拟细胞随时间演变的行为,如细胞周期、分化过程或对外界刺激的适应性响应。
  3. 预测能力有限: 由于未能整合多层次信息,传统模型在预测细胞对未曾见过的扰动(如新药作用)的响应时,其准确性和鲁棒性往往不足。

相比之下,虚拟细胞模型展现出显著的系统性优势,主要体现在其“多维度”和“多状态”模拟能力:

  1. 多维度整合: 虚拟细胞能够无缝整合从分子到细胞再到组织层面的信息,实现跨尺度的建模与模拟 89。这意味着它不仅能考虑单个基因或蛋白质的功能,还能模拟这些分子如何组装成信号通路、代谢网络,进而影响细胞的整体行为(如增殖、凋亡、迁移),甚至细胞间的相互作用,最终影响组织层面的功能 1011。例如,一个虚拟细胞模型可以模拟肿瘤细胞内部的基因突变如何导致其增殖失控,并进一步模拟这些肿瘤细胞如何与周围的免疫细胞、基质细胞相互作用,形成复杂的肿瘤微环境。这种全面的视角有助于揭示复杂生物学现象的深层机制。
  2. 多状态模拟: 虚拟细胞能够模拟细胞在不同生理和病理状态下的行为,包括正常生理状态、疾病状态(如感染、癌症、神经退行性疾病)以及受到干预(如药物治疗、基因编辑)后的响应 112。通过在虚拟环境中模拟这些状态的转变,科学家可以系统性地探索疾病发生发展的机制,评估不同治疗策略的潜在效果,甚至设计个性化的治疗方案。例如,在药物研发中,虚拟细胞可以模拟健康细胞和癌细胞对某种药物的不同反应,从而预测药物的疗效和潜在副作用。在神经退行性疾病研究中,虚拟细胞可以模拟神经细胞在应激条件下的损伤过程,并测试保护性干预措施的效果。

总而言之,虚拟细胞通过其强大的数据整合能力和复杂的计算模拟技术,克服了传统还原论模型在系统性、动态性和预测能力方面的局限,为理解和干预生命过程提供了前所未有的工具和视角 1。

2. 虚拟细胞的构建技术路径

2.1 数据基础:多源异构生物数据整合

虚拟细胞的构建,其核心在于对海量、多源、异构生物学数据的有效整合与深度挖掘。这些数据如同构成细胞生命的“信息砖块”,为构建高保真度的虚拟模型提供了基础。其中,单细胞测序、空间转录组、蛋白质组等多种组学数据扮演着至关重要的角色。

  1. 单细胞测序数据: 单细胞RNA测序(scRNA-seq)技术能够揭示单个细胞层面的基因表达谱,从而识别细胞亚型、追踪细胞分化路径、理解细胞异质性。例如,通过单细胞测序,研究人员可以发现以往未知的细胞群,如骨关节炎(OA)中炎症性软骨细胞和炎症前软骨细胞的新亚群 13。这些高分辨率的细胞表型数据是虚拟细胞模型能够区分不同细胞状态、模拟细胞命运决定的关键。除了RNA测序,单细胞多组学技术能够进一步提供更全面的细胞信息,例如对氧诱导视网膜病变模型中Müller细胞的异质性研究就使用了单细胞多组学分析方法 14。

  2. 空间转录组数据: 空间转录组学技术则弥补了单细胞测序丢失空间信息的不足,它能够在保留组织结构完整性的前提下,测量细胞在原位(in situ)的基因表达模式 13。这对于理解细胞间的相互作用、细胞微环境以及组织层面的功能组织至关重要。例如,通过空间转录组数据,可以确定特定疾病相关基因在组织中的精确位置,如骨关节炎中大多数OA相关差异表达基因位于关节表面和浅表区域 13。一些方法甚至能利用未匹配的单细胞RNA-seq数据来预测空间转录组数据的细胞组成,从而更精准地绘制细胞图谱 15。虚拟细胞模型通过整合空间信息,能够模拟细胞在真实组织环境中的行为,例如肿瘤细胞与周围免疫细胞的局部相互作用。

  3. 蛋白质组学数据: 蛋白质是细胞功能的主要执行者,蛋白质组学研究细胞内所有蛋白质的表达、修饰和相互作用。这些数据对于理解细胞的信号通路、代谢网络以及细胞骨架等结构性功能至关重要。图像引导的单细胞多组学数据(如DNA和RNA FISH结合核蛋白组学)能够深入揭示基因调控与染色质、信使RNA和关键蛋白空间位置之间的关系,甚至可以用于构建虚拟细胞核模型 16。

数据标准化与质量控制的关键问题:

虚拟细胞模型需要整合来自不同实验室、不同平台、不同批次的异构数据,这带来了严峻的数据标准化和质量控制挑战。

  1. 数据标准化: 不同组学数据类型(如RNA计数、蛋白丰度、基因组变异)本身具有不同的数据格式、测量单位和噪音水平。在整合之前,必须进行严格的标准化处理,以消除批次效应、实验误差和技术差异,确保数据可比性。这包括数据归一化、批次校正以及将不同数据类型映射到统一的表征空间。

  2. 质量控制(QC): 数据的质量直接决定了虚拟细胞模型的准确性和可靠性。质量控制贯穿于数据获取、处理和分析的整个流程。它包括对原始数据的过滤(如去除低质量细胞、死细胞)、异常值检测、缺失值处理以及验证数据的一致性和完整性。对于整合大规模、多来源的数据集,需要开发先进的质量管理协议来防止错误并识别和纠正现有错误,因为即使有全面的质量保证协议,数据错误率仍可能很高 17。

通过有效的多源异构生物数据整合,并辅以严格的数据标准化和质量控制流程,虚拟细胞模型得以建立一个坚实的数据基石,从而为后续的AI与生物物理模型融合提供高质量的输入,实现对细胞生命活动的精确模拟。

2.2 核心技术:AI与生物物理模型的融合

虚拟细胞的构建不仅仅是数据的堆砌,更关键在于将这些多尺度、多模态的生物数据转化为具有预测能力的动态模型。这需要融合多种核心技术,其中人工智能(AI)和生物物理模型扮演着至关重要的角色,通过协同作用构建虚拟细胞的多维度行为预测能力。

  1. 生成式AI(如大神经网络)的作用:
    近年来,以大型神经网络为代表的生成式AI在处理复杂数据、发现隐藏模式方面展现出强大能力,为虚拟细胞的构建提供了前所未有的机遇 118。

    • 学习复杂生物模式: 生成式AI,特别是大型语言模型(LLMs)和生成模型,能够学习和捕捉来自海量生物数据集的复杂模式,例如基因调控网络、蛋白质相互作用以及细胞对不同刺激的响应。它们可以从细胞图谱中预训练获得丰富的细胞表征,并通过在新的生物学任务中进行微调来适应特定需求 19。
    • 生成式模拟与预测: AI模型能够通过学习细胞的动态行为,生成在不同条件下的细胞状态预测。例如,利用AI可以预测药物对细胞响应的影响,或在基因编辑后细胞行为的变化 2021。这种能力使得虚拟细胞可以在计算环境中进行“虚拟实验”,显著加速药物发现和生物医学研究 120。
    • 数据整合与表征: AI能够将不同来源、不同尺度的异构生物数据(如基因组、转录组、蛋白质组数据)整合到一个统一的数学框架中,并学习出高维度的细胞嵌入(cell embeddings),有效表征细胞的内部活动和状态 19。这有助于克服传统方法在处理多模态数据时的局限性。
    • 可编程的细胞表型控制: 结合AI代理(AI agents),虚拟细胞甚至有望实现对细胞表型的可编程控制和细胞回路的设计 22。
  2. 生物物理建模(如细胞力学模拟)的贡献:
    尽管AI在模式识别和预测方面表现出色,但生物系统本质上受物理定律的支配。生物物理模型从第一性原理出发,基于已知的物理化学规律(如扩散、分子动力学、力学)来描述细胞行为,为虚拟细胞提供了坚实的物理基础和可解释性 23。

    • 捕捉物理过程: 生物物理模型能够精确模拟细胞内的物理过程,例如分子的扩散 24、化学反应动力学、细胞膜的流动性、细胞骨架的机械力学特性以及细胞如何感知和响应机械信号 2023。例如,在Virtual Cell软件中,可以使用基于规则的模型来描述细胞内复杂的反应网络,并利用各种ODE或随机求解器进行模拟 2526。
    • 空间和时间尺度: 细胞的形状、亚细胞区室的大小以及分子在细胞质中的空间分布都会影响分子间的相互作用,进而影响细胞行为 24。生物物理模型,特别是偏微分方程(PDE)和基于个体(agent-based)的模型,能够纳入细胞的几何结构和空间维度,模拟分子在不同亚细胞区域的分布和运动,以及细胞之间的力学相互作用 2324。
    • 可解释性与机制洞察: 与AI的“黑箱”特性不同,生物物理模型通常具有更好的可解释性,它们明确地表达了生物学过程背后的物理机制,有助于研究人员理解“为什么”细胞会表现出某种行为。例如,通过模拟细胞骨架蛋白在协调细胞行为中的机械作用,可以深入理解细胞如何利用物理信号来执行功能 20。
  3. 动态交互算法(如基因调控网络)的协同:
    细胞是一个高度动态和交互的系统,基因调控网络、信号转导通路以及细胞间的通讯是其核心。动态交互算法旨在捕捉这些复杂、非线性的相互作用。

    • 基因调控网络建模: 通过分析基因表达数据,可以构建基因调控网络模型,描述基因之间如何相互作用,从而控制细胞命运和功能 2728。这些网络通常采用微分方程、布尔网络或基于代理的模型来表示,以模拟基因表达随时间的变化。
    • 信号转导通路模拟: 细胞通过复杂的信号转导通路响应外界刺激,动态交互算法可以模拟这些通路中蛋白质的磷酸化、激活和失活过程,从而预测细胞对激素、生长因子或药物的响应。
    • 细胞间通讯: 在多细胞系统中,细胞通过直接接触、分泌分子或纳米通讯(如FRET)进行交流。动态交互算法能够模拟这些细胞间的通讯方式如何影响群体行为,例如组织形成、免疫响应或肿瘤微环境中的相互作用 29。
    • 融合AI与生物物理: 动态交互算法常常作为AI和生物物理模型之间的桥梁。例如,AI可以从数据中学习基因调控网络的参数,然后将这些参数整合到基于ODE或随机过程的生物物理模型中进行动态模拟。反之,生物物理模型模拟的细胞力学反馈也可以作为AI模型的输入,指导细胞行为的进一步预测。例如,有研究提出将生成式AI与时空多尺度生物物理建模相结合,以构建虚拟细胞,模拟和评估药物和基因编辑技术对各种细胞动态行为的治疗效果 20。

综上所述,虚拟细胞的构建是一个多学科交叉的工程。生成式AI提供强大的模式识别和预测能力,能够处理海量复杂数据并进行高维表征;生物物理模型确保了模型的物理真实性和可解释性;而动态交互算法则负责捕捉细胞内部和细胞间复杂的动态反馈和调节机制。这三者的深度融合,使得虚拟细胞能够实现对生命现象的多维度、高精度和动态预测,为未来的生物医学研究和临床应用奠定基础。

2.3 验证机制:从计算模拟到实验闭环

虚拟细胞模型虽然具备强大的预测能力,但其科学价值和可靠性最终需要通过严格的实验验证来确立。这一过程形成一个“计算模拟-实验验证-模型迭代优化”的闭环,确保虚拟细胞能够准确反映真实生物系统的行为 30。

虚拟细胞模型的验证流程通常包括以下几个关键环节:

  1. 预测生成与假设建立: 基于构建完成的虚拟细胞模型,研究人员可以运行模拟,对细胞在特定条件(例如药物处理、基因扰动、环境变化)下的行为作出预测。这些预测随后被转化为具体的生物学假设,用于指导实验设计。例如,虚拟细胞模型可能会预测某个基因的敲除将导致细胞增殖速率显著下降,或某种药物会特异性抑制某一信号通路。

  2. 实验验证方法: 为了检验虚拟细胞模型的预测,需要采用多种实验技术在真实的生物系统中进行验证。

    • CRISPR扰动实验: 基因编辑技术,特别是CRISPR-Cas9系统,是验证虚拟细胞模型预测的关键工具。研究人员可以根据模型的预测,对特定基因进行敲除(loss-of-function)、过表达(gain-of-function)或条件性编辑,然后观察细胞表型是否与模型预测一致 3132。例如,如果虚拟细胞模型预测抑制SLC7A1会影响肿瘤细胞对铁死亡的敏感性,实验可以通过CRISPR技术敲除SLC7A1基因,然后观察肿瘤细胞在铁死亡诱导剂作用下的存活率。基因筛选(Genetic Screens),特别是高通量CRISPR筛选,能够系统性地评估基因功能,为虚拟细胞模型提供大规模的验证数据 3132。
    • 类器官平台测试: 类器官(organoids)作为一种体外三维细胞培养模型,能够模拟真实器官的复杂结构和功能,为虚拟细胞模型的验证提供了更接近生理环境的平台 33。例如,人类骨髓类器官已被用于血液恶性肿瘤的疾病建模和治疗靶点验证 33。虚拟细胞模型可以预测药物在类器官中的效果,然后通过在类器官上进行药物筛选来验证这些预测 34。这种平台允许研究人员在受控的实验条件下,测试模型对细胞分化、组织形成、药物响应以及细胞间相互作用的预测能力。
    • 其他体外/体内实验: 除了CRISPR和类器官,传统的细胞生物学实验(如细胞增殖、迁移、凋亡测定)、生化分析(如蛋白质表达、酶活性检测)、以及更复杂的体内动物模型,都可以用来验证虚拟细胞模型的预测。例如,虚拟现实(VR)技术也在神经科学研究中被用来进行闭环反馈的感官刺激和行为学实验,为模拟神经系统行为的虚拟模型提供验证手段 3536。
  3. 模型迭代优化与闭环学习: 实验验证的结果至关重要,它决定了虚拟细胞模型的下一步发展方向。

    • 评估预测准确性: 比较实验结果与模型预测之间的符合程度。如果预测与实验结果一致,则增强了模型的置信度;如果不一致,则表明模型存在局限性或需要修正。
    • 参数校准与结构调整: 对于不一致的情况,研究人员需要分析原因,并对虚拟细胞模型进行迭代优化。这可能包括:
      • 参数校准: 调整模型中一些不确定或经验性参数的数值,使其更符合实验观测。这常常涉及优化算法,例如层次参数优化技术 37。
      • 模型结构调整: 如果简单的参数调整不足以解决问题,可能需要重新审视模型的底层假设,引入新的生物学机制、更新基因调控网络、或改进物理模型(例如细胞力学模拟)的方程。
      • 整合新的数据: 将新的实验数据整合到模型中,进一步丰富模型的信息含量,提升其表征能力。
    • 闭环学习: 理想的虚拟细胞模型应具备从实验数据中“学习”的能力,实现自动化或半自动化的迭代优化。例如,有研究提出了一个“闭环”框架,通过在模型微调过程中整合扰动数据,显著提高了T细胞激活预测的准确性,并识别出新的治疗靶点和通路,这标志着虚拟细胞模型迈向“学习型”和“自适应型”的关键一步 38。

通过这种严格的验证与迭代优化机制,虚拟细胞模型能够不断地逼近真实生物系统的复杂性,提高其在药物研发、疾病机制研究和基础生物学领域的应用价值。

3. 虚拟细胞的典型应用场景

3.1 药物研发与精准医疗

虚拟细胞技术在药物研发和精准医疗领域展现出巨大的应用潜力,能够显著加速新药发现、优化药物筛选过程、评估药物毒性,并最终实现个性化用药方案设计。传统药物研发通常耗时长、成本高且失败率高,而虚拟细胞通过在计算环境中模拟细胞对药物的响应,为这些挑战提供了创新的解决方案。

  1. 药物响应预测与筛选:
    虚拟细胞能够模拟药物分子与细胞内靶点相互作用的动态过程,进而预测细胞对药物的反应。这包括药物的吸收、分布、代谢、排泄(ADME)过程,以及药物在细胞内部引发的信号通路级联反应。例如,通过构建癌细胞的虚拟模型,研究人员可以模拟不同抗癌药物对肿瘤细胞增殖、凋亡或分化的影响。这使得科学家能够在进入昂贵的湿实验之前,大规模地筛选潜在的药物分子,显著提高药物发现的效率和成功率。

    • SLC7A1抑制剂筛选的案例: 在骨肉瘤的研究中,溶质载体家族转运蛋白SLC7A1被发现与肿瘤的恶性进展和不良预后高度相关。单细胞分析显示,SLC7A1在骨肉瘤细胞中表达突出,并与多种恶性肿瘤特征相关,还能调控肿瘤细胞与巨噬细胞的相互作用以及巨噬细胞的极化。通过虚拟筛选和细胞热转移分析(CETSA),研究人员已成功鉴定出Cepharanthine等潜在的SLC7A1抑制剂39。虚拟细胞模型可以在更广泛的药物库中,进一步模拟这些抑制剂与SLC7A1的结合效率、细胞内作用机制以及对骨肉瘤细胞生长和转移的抑制效果,从而加速针对SLC7A1的靶向药物开发。
    • 高通量药物筛选: 基于基因组学和CRISPR-Cas9筛选的药物靶点优先级排序,能够识别出癌症治疗的潜在靶点4041。虚拟细胞可以整合这些高通量筛选数据,构建包含多个潜在靶点的细胞模型,并模拟药物对这些靶点的组合作用,从而预测药物联用的效果,并筛选出协同增效的药物组合。
    • AI赋能药物发现: 人工智能,特别是深度生成模型,能够预测细胞对新型化学扰动的转录反应,从而实现在硅(in-silico)药物筛选,并已成功推荐小细胞肺癌和结直肠癌的候选药物42。这种AI驱动的虚拟细胞框架,可以极大地扩展药物筛选的范围和速度,为创新药物开发提供了新范式43。
  2. 毒性评估:
    药物的安全性是其能否上市的关键。虚拟细胞可以用于预测药物对非靶细胞或正常组织的潜在毒性作用。通过模拟药物在不同细胞类型中的代谢途径和作用机制,虚拟细胞能够识别潜在的脱靶效应和细胞损伤机制。例如,虚拟细胞可以模拟肝细胞对药物的代谢过程以及产生的毒性产物对细胞功能的影响。有研究提出,虚拟细胞基于的检测方法(Virtual Cell Based Assay, VCBA)可以用于化学品风险评估,模拟化学物质在体外测试系统中的动态过程,从而加速化学品安全评估44。

  3. 个性化用药方案设计:
    精准医疗的核心是根据患者的个体特征(如基因组信息、疾病亚型、生物标志物)来制定最合适的治疗方案。虚拟细胞通过整合患者特异性的组学数据(如肿瘤患者的基因突变谱),构建个体化的虚拟细胞模型。

    • 肿瘤治疗: 对于肿瘤患者,虚拟细胞可以模拟其特异性肿瘤细胞对不同化疗药物、靶向药物或免疫疗法的反应。例如,通过模拟患者肿瘤细胞的突变状态和信号通路活性,虚拟细胞可以预测哪些药物对该患者最有效,同时规避耐药性或严重的副作用。这有助于医生为患者选择最佳的治疗方案,避免无效治疗和不必要的毒性。
    • 神经系统疾病: 虚拟细胞在神经退行性疾病的个性化治疗中也显示出潜力。例如,通过构建特定神经元亚型的虚拟细胞模型,可以模拟神经元功能障碍,并评估针对这些病理机制的个性化治疗策略。研究表明,在神经退行性疾病中,通过计算模型重建轴突内线粒体动态,可模拟Dst基因敲除小鼠中线粒体运输异常和结构变形,并预测敲除Nefl可缓解神经退行性进展45。
    • 罕见病和特殊人群: 对于罕见病患者或儿童等特殊人群,由于临床数据有限,个性化用药方案的设计尤为困难。虚拟细胞可以基于有限的患者数据进行建模和预测,为这些患者提供更精准的治疗指导。

综上所述,虚拟细胞正从根本上改变药物研发的范式,从传统的试错法转向基于机制理解和预测的理性设计。它有望大幅降低研发成本和时间,提高药物开发的成功率,并最终实现更加精准和个性化的医疗服务。

3.2 疾病机制与病理研究

虚拟细胞在深入理解复杂疾病的发生、发展机制及其病理生理过程方面发挥着不可替代的作用。通过在计算环境中模拟疾病状态下的细胞行为和相互作用,虚拟细胞能够揭示传统实验方法难以捕捉的动态过程和潜在机制,为疾病诊断和治疗提供新的理论基础。

  1. 肿瘤微环境互作的解析:
    肿瘤并非由单一的癌细胞组成,而是由癌细胞、免疫细胞、成纤维细胞、血管细胞以及细胞外基质等多种组分构成的复杂微环境(Tumor Microenvironment, TME)。虚拟细胞能够模拟这些不同细胞类型之间的动态相互作用,从而揭示肿瘤生长、侵袭、转移以及对治疗响应的关键机制。

    • 细胞-免疫微环境交互: 免疫系统在肿瘤的发生发展中扮演着双重角色,既能抑制肿瘤,也可能被肿瘤利用促进其生长。虚拟细胞可以通过构建包含肿瘤细胞和免疫细胞(如T细胞、巨噬细胞)的模型,模拟它们之间的动态相互作用。例如,研究发现细胞-T细胞结合的动力学模型可以改变免疫控制癌症的时间尺度,这表明在数学模型中明确包含结合区室对于理解肿瘤控制和进展的动力学至关重要 46。通过调整虚拟模型中的参数,研究人员可以探索不同免疫细胞亚型在肿瘤微环境中的招募、激活、浸润以及对肿瘤细胞的杀伤效率。这有助于理解肿瘤免疫逃逸的机制,并为开发新型免疫疗法提供靶点。
    • 基质成分的作用: 肿瘤微环境中的细胞外基质(Extracellular Matrix, ECM),特别是胶原纤维的排列,被认为是免疫逃逸的机制之一。计算模拟框架可以区分免疫细胞平行或垂直于胶原纤维方向迁移的两种假设,揭示胶原模式如何为免疫提供保护,从而导致疾病晚期免疫覆盖率降低 47。虚拟细胞模型可以整合这些空间数据,模拟ECM如何影响细胞的迁移、增殖和药物渗透,从而更全面地理解肿瘤进展。
  2. 神经退行性疾病的研究:
    神经退行性疾病(如阿尔茨海默病、帕金森病)的病理机制复杂多样,涉及神经炎症、蛋白质错误折叠、线粒体功能障碍、神经元死亡等多个环节 48。虚拟细胞模型在理解这些疾病的共同机制和探索潜在治疗靶点方面具有独特优势。

    • 神经炎症的模拟: 神经炎症是许多神经退行性疾病的核心特征。虚拟细胞可以模拟小胶质细胞和星形胶质细胞等神经免疫细胞在疾病进程中的激活、炎症介质的释放以及对神经元的影响。例如,在阿尔茨海默病(AD)研究中,虚拟筛选已成功识别出脑渗透性的半乳糖凝集素-3(Gal-3)抑制剂FJMU1887,该抑制剂能有效抑制BV-2小胶质细胞的炎症反应和TNF-α的产生,并改善AD小鼠模型的认知功能 49。虚拟细胞模型可以进一步模拟Gal-3在AD病理中的作用,以及FJMU1887如何通过干扰Gal-3-TREM2相互作用来减轻神经炎症和淀粉样β蛋白负担。
    • 共同机制的探索: 神经退行性疾病往往存在共同的细胞和分子事件,例如氧化应激、线粒体功能障碍和神经血管相互作用 48。虚拟细胞能够整合来自不同疾病模型的组学数据和生理数据,构建统一的神经元损伤模型,从而识别跨疾病类别的通用治疗靶点。例如,通过模拟线粒体动态和功能障碍,虚拟细胞可以预测针对线粒体保护的干预措施对神经元存活和功能的影响。
    • ZBP1在心肌缺血再灌注损伤中的作用: ZBP1是一种Z-DNA结合蛋白,在心肌缺血再灌注(I/R)损伤中发挥关键作用,通过诱导心肌细胞的PANoptosis(一种程序性细胞死亡形式)加重心肌损伤。虚拟筛选已发现小分子化合物MSB能够高亲和力结合ZBP1,并在体外和体内有效减轻心肌I/R损伤 50。虚拟细胞模型可以模拟ZBP1在缺血再灌注条件下激活PANoptosome复合物(ZBP1/RIPK3/CAS8/CAS6)的动态过程,并评估MSB如何通过阻断ZBP1来抑制心肌细胞PANoptosis,从而为心肌I/R损伤的治疗提供新的策略。
  3. 其他疾病机制的揭示:
    虚拟细胞的应用范围远不止肿瘤和神经退行性疾病。例如,在酒精性肝病(ALD)的研究中,FOXO3被认为是与长寿相关的转录因子,通过促进抗氧化应激反应和抑制炎症来发挥作用。虚拟筛选已识别出FOXO3激动剂214991,并在动物模型中证实其能显著减轻酒精诱导的肝损伤 51。虚拟细胞模型可以模拟FOXO3在肝细胞中的调控网络以及214991如何通过激活FOXO3通路来发挥抗炎、抗氧化和抗凋亡作用。此外,虚拟细胞也被用于研究败血症(Sepsis)的病理生理学,通过结合类器官、组学、免疫细胞表型分析和建模来识别新的治疗靶点,例如理解中性粒细胞-内皮细胞相互作用在败血症进展中的关键作用 52。

通过这些深入的疾病机制研究,虚拟细胞不仅能增进我们对复杂疾病的理解,还能为开发更有效、更精准的治疗方案指明方向。

3.3 基础生物学与发育研究

虚拟细胞技术在基础生物学和发育生物学领域也展现出巨大的潜力,它提供了一个强大的平台来模拟和理解细胞在发育、稳态和病理过程中的动态行为、命运决定以及组织形成机制。通过将复杂的生物学过程分解为可计算的模型,虚拟细胞有助于揭示细胞功能和组织形成的系统性原理。

  1. 细胞动态行为的模拟:
    虚拟细胞能够模拟细胞的各种动态行为,如增殖、迁移、分化、凋亡以及形状变化等。这些行为是生物体正常发育和功能维持的基础。

    • 上皮-间质转化 (EMT) 模拟: 上皮-间质转化(EMT)是一种细胞分化过程,使上皮细胞获得间充质细胞的特性,包括增加的迁移能力和侵袭性。EMT在胚胎发育、伤口愈合、干细胞行为中至关重要,但也病理性地参与纤维化和癌症进展 53。虚拟细胞模型可以精确模拟EMT过程中的基因表达重编程和非转录变化,以及TGFβ等信号通路如何启动和控制这一过程 5354。通过模拟EMT,虚拟细胞可以帮助科学家理解细胞如何从一个表型转变为另一个表型,以及这种可塑性如何影响疾病进程,例如肿瘤转移 5556。这些模型甚至可以区分不同EMT转录因子(如SNAIL, ZEB)的非冗余功能,从而解释EMT在不同疾病中作用的争议 57。
    • 细胞迁移与相互作用: 细胞的迁移和相互作用是组织形成、免疫应答和疾病进展的关键。例如,在胚胎发育过程中,细胞通过迁移形成复杂的组织结构;在肿瘤中,癌细胞通过迁移扩散到身体其他部位。虚拟细胞可以通过基于代理(agent-based)的模型来模拟单个细胞的运动、细胞间的粘附、排斥以及与其他细胞或细胞外基质的相互作用,从而预测细胞群体的集体行为,例如在淋巴生发中心,B细胞的突变和选择过程可以被代理模型有效模拟 58。
  2. 早期胚胎发育模式模拟(如多能干细胞分化):
    早期胚胎发育是一个高度协调和精密的细胞分化和模式形成过程,涉及多能干细胞(PSC)向不同细胞谱系的特化。虚拟细胞为理解这一复杂过程提供了独特的视角。

    • 多能干细胞命运决定: 人类多能干细胞(hPSC)在生物医学研究和再生医学中具有巨大潜力,但其扩增和分化过程的精确控制仍面临挑战 59。虚拟细胞模型可以整合基因调控网络和信号通路数据,模拟多能干细胞(如胚胎干细胞,ES)如何通过调控关键转录因子(如Oct4和Nanog)的表达模式来维持多能性或启动分化 60。这些模型能够揭示细胞命运决定的机制,例如双稳态开关如何协同调控ES细胞向胚外滋养外胚层和原始内胚层谱系分化 60。通过模拟不同的生长因子或化学诱导剂对干细胞分化的影响,虚拟细胞可以帮助优化体外干细胞培养方案,以高效获得特定细胞类型。
    • 组织和器官形成: 虚拟细胞可以扩展到模拟多细胞系统,研究细胞如何自组织形成复杂的组织和器官结构。这涉及到细胞增殖、凋亡、分化、迁移以及细胞间通讯的协调。例如,计算模型可以模拟再生过程中的组织模式形成,帮助理解器官和组织的再生机制 61。此外,类器官(organoids)作为体外培养的三维结构,能够重现体内器官的关键特征,它们与虚拟细胞模型结合,可以用于研究器官发育、疾病建模和药物筛选 62。
    • 系统性理解细胞功能: 虚拟细胞通过整合多尺度数据和复杂的计算模型,能够提供对细胞功能和组织形成的系统性理解。它不仅关注单个基因或蛋白质的作用,更强调这些组分如何在动态网络中相互作用,共同决定细胞的整体行为和宏观表型。例如,通过模拟细胞系谱追踪(cell lineage tracing),虚拟细胞可以重建细胞从祖细胞到终末分化细胞的整个发育轨迹,揭示基因调控如何指导细胞命运决定 6364。这有助于发现新的发育调控因子,并为再生医学提供理论指导。

通过这些在基础生物学和发育研究中的应用,虚拟细胞不仅加深了我们对生命基本过程的理解,也为解决发育缺陷、组织损伤和衰老相关问题提供了新的思路。

4. 挑战与未来发展方向

4.1 技术挑战与瓶颈

尽管虚拟细胞在生物医学领域展现出巨大的应用潜力,但在其全面发展和实际应用过程中,仍面临一系列显著的技术挑战和瓶颈。这些挑战主要体现在数据整合、模型可解释性以及生物真实性验证等关键环节。

  1. 数据整合的复杂性(如跨尺度数据关联):
    虚拟细胞的核心在于整合海量、多源、异构的生物数据,以构建一个全面的细胞模型。然而,要实现这一目标,数据整合本身就是一项艰巨的任务。

    • 异构性与标准化: 生物数据来自不同的实验平台和技术(如单细胞测序、蛋白质组学、代谢组学、表观遗传学、成像数据等),它们具有不同的数据格式、测量尺度、信噪比和实验偏差。将这些异构数据有效整合到一个统一的框架中,需要开发先进的数据标准化、归一化和批次效应校正方法。例如,整合基因表达数据和蛋白质丰度数据时,需要解决两者之间复杂的非线性关系和时间滞后问题。
    • 跨尺度关联: 虚拟细胞模型需要涵盖从分子(如基因、蛋白质)到亚细胞器(如线粒体、细胞核)、细胞(如细胞类型、状态)再到组织(如细胞微环境、器官结构)的多个生物学尺度。这意味着模型不仅要描述各个尺度内部的生物学过程,更要建立跨尺度之间的精确关联机制。例如,如何将某个基因的微小变化与细胞层面上的增殖、迁移行为,乃至组织层面的病理生理变化精确地联系起来,仍是一个巨大的挑战。现有的数据往往在不同尺度上存在信息鸿沟,缺乏能够有效连接这些尺度的实验数据或计算方法。
    • 动态性与时序性: 细胞是一个动态变化的系统,其状态和行为是随时间演变的。传统的静态数据整合难以捕捉这种动态过程。因此,如何整合时间序列数据,并建立能够反映细胞动态变化的因果关系模型,是虚拟细胞模型提高预测能力的关键。
  2. 模型可解释性(如AI黑箱问题):
    随着人工智能(AI),特别是深度学习技术在虚拟细胞构建中的广泛应用,模型的可解释性问题日益突出。

    • “黑箱”特性: 大规模神经网络模型虽然在预测精度上表现卓越,但其内部运作机制往往不透明,被称为“黑箱”模型。研究人员难以直接理解模型为何做出特定预测,哪些输入特征对输出结果贡献最大,以及模型是否捕捉到了真实的生物学因果关系,而非仅仅是统计关联。这种“黑箱”特性极大地限制了虚拟细胞模型在发现新生物学机制方面的能力。
    • 信任与验证: 在生物医学领域,模型的决策往往涉及到重要的临床判断。如果模型缺乏可解释性,医生和研究人员就难以信任其预测结果,也无法根据模型建议进行进一步的实验或临床干预。例如,当虚拟细胞预测某种药物对患者有效时,医生需要了解模型作出此判断的生物学依据。
    • 因果推断的挑战: 生物学研究的核心目标是揭示因果关系,而AI模型在区分因果与相关性方面仍面临挑战。缺乏可解释性使得从AI驱动的虚拟细胞模型中提取可靠的因果推断变得困难。
  3. 生物真实性验证(如与真实实验的匹配度):
    虚拟细胞模型最终的价值在于其对真实生物系统行为的准确模拟和预测。因此,模型与真实实验结果的匹配度是衡量其成功与否的关键指标,但这一验证过程本身也充满挑战。

    • 实验条件的差异: 虚拟细胞模型是在理想化的计算环境中运行的,而真实实验(无论是体外细胞培养、类器官还是体内动物模型)都受到无数复杂且难以完全控制的因素影响,例如批次效应、环境条件波动、个体差异等。这使得模型预测与实验结果之间可能存在偏差。
    • 复杂系统与高维度数据: 细胞系统极其复杂,涉及的变量和相互作用是海量的。尽管模型试图捕捉这种复杂性,但任何模型都只能是对真实世界的简化。如何设计足够多的、具有代表性的实验来全面验证模型在高维度空间中的预测能力,是一个巨大的挑战。
    • 验证标准的缺失: 目前,对于如何定义虚拟细胞模型的“成功”或“高保真”,尚缺乏统一的、被广泛接受的验证标准和基准数据集。例如,模型预测的某个基因表达水平与实验测量值相差多少才能被认为是“足够准确”?模型预测的细胞迁移速度与实验观察值之间多大的差异可以接受?建立类似“图灵测试”的评估框架,以客观衡量虚拟细胞的生物真实性,是亟待解决的问题 65。
    • 动态验证的困难: 验证模型对细胞动态过程(如细胞分化轨迹、疾病进展)的预测更为困难,因为它需要在多个时间点和状态下进行对比,对实验数据的要求更高。

克服这些技术挑战和瓶颈,需要跨学科的深度合作,持续的技术创新,以及开放科学社区的共同努力,才能真正推动虚拟细胞从概念走向成熟,并最终赋能生物医学研究和临床应用。

4.2 跨学科协作与开放科学

虚拟细胞的构建与发展是一个极其复杂的系统工程,它超越了单一学科的范畴,天然地要求生物学家、计算科学家、临床医生等多领域专家进行深度、持续的协作。同时,开放科学的理念和实践,如公共数据共享平台和标准化建模工具的推广,对加速虚拟细胞技术从实验室走向广泛应用至关重要。

  1. 多领域协作的必要性:

    • 生物学家的核心作用: 生物学家是虚拟细胞的“需求方”和“验证方”。他们对细胞生物学机制、疾病病理、实验方法和数据解读拥有深入的专业知识。生物学家需要提出明确的科学问题,提供高质量的实验数据,并基于其生物学直觉和经验对模型的预测结果进行批判性评估。他们是确保虚拟细胞模型生物真实性和应用价值的基石。
    • 计算科学家的技术支撑: 计算科学家(包括计算机科学、数学、物理学、生物信息学专家)是虚拟细胞的“建造者”和“工具提供者”。他们负责开发和实现先进的计算模型(如AI算法、生物物理模拟、统计推断),处理和整合海量异构数据,以及设计高效的模拟和分析框架。没有计算科学家的专业技能,虚拟细胞的复杂建模和大规模数据处理将无法实现。
    • 临床医生的转化桥梁: 临床医生是虚拟细胞走向临床应用的“引导者”。他们熟悉疾病的临床表现、诊断标准、治疗方案以及患者的个体差异。临床医生能够提供宝贵的临床数据和反馈,帮助虚拟细胞模型聚焦于临床亟待解决的问题,并评估模型在疾病诊断、预后判断和治疗决策方面的实际价值和可操作性。例如,数学模型与深度学习框架结合,可以生成个性化的自适应治疗方案,并在虚拟患者模型中表现出优于临床标准治疗的效果,为临床实践提供了新的思路 66。
    • 学科交叉的融合效应: 虚拟细胞的成功依赖于这些不同领域的知识和方法的有机融合。生物学家提供生物学洞察,计算科学家将这些洞察转化为可计算的模型,临床医生则将模型与现实世界的医疗需求相结合。例如,一个成功的虚拟细胞模型可能需要生物学家提出细胞器互作的假说,计算科学家利用多尺度模拟和深度学习来构建模型,再由临床医生通过患者数据来验证模型在预测药物反应方面的有效性。这种深度融合能够催生出单一学科难以实现的创新突破。
  2. 开放科学对推动虚拟细胞发展的作用:
    开放科学倡导开放获取、数据共享和协作研究,这与虚拟细胞的发展需求高度契合。

    • 公共数据共享平台: 虚拟细胞的构建需要海量的多模态生物数据。建立和推广公共数据共享平台(如中国国家基因库CNGBdb 67、PhenoDB 68、各种“组学”数据库)对于汇集全球研究力量、加速模型开发至关重要。这些平台不仅提供原始数据,还应包含标准化处理后的数据、元数据以及质量控制报告,确保数据的可用性和可信度。数据共享能够避免重复实验,降低研究成本,并让更多研究者能够访问和利用数据来构建和验证虚拟细胞模型。
    • 标准化建模工具与框架: 虚拟细胞模型的复杂性意味着需要统一的建模语言、数据格式和软件工具,以促进模型的可重复性、互操作性和社区协作。开发和推广开源、标准化的建模框架(如生物物理模拟库、AI模型开发工具、可视化平台)能够降低入门门槛,使得不同实验室能够共享和修改彼此的模型,从而加速模型的迭代优化。例如,专注于病毒共识基因组重建的ViReflow管线通过利用亚马逊网络服务(AWS)云计算资源和Reflow系统,实现了病毒序列数据集的快速分析,这体现了标准化工具在促进大规模生物数据分析方面的潜力 69。
    • 促进模型验证与基准测试: 开放科学精神鼓励研究者分享他们的模型代码和验证数据,从而促进同行评议和基准测试。通过共享模型,不同团队可以独立地测试模型的性能和鲁棒性,发现模型的优点和局限性,并推动建立一套公认的模型验证标准,类似于“如何在细胞中构建虚拟细胞”的愿景,强调了数据需求、评估策略和社区标准在确保生物准确性和广泛实用性方面的作用 18。
    • 加速知识传播与创新: 开放科学能够促进研究成果的快速传播和知识共享,有助于跨学科领域的研究人员及时了解最新的技术进展和研究范式。这种信息流动对于激发创新思想、形成新的研究合作至关重要。例如,通过机器学习算法发现抗衰老药物的案例表明,即使是小型和异构的药物筛选数据,AI也能最大化其利用价值,这为开放科学方法在早期药物发现中的应用铺平了道路 70。

总而言之,虚拟细胞的未来发展将紧密依赖于构建一个强大的多学科协作网络和实践开放科学的社区。这不仅能解决技术难题,还能确保虚拟细胞模型能够真正地服务于生物学发现、疾病理解和临床实践。

4.3 临床转化与产业应用前景

虚拟细胞技术不仅是科研领域的强大工具,其在临床转化和产业应用方面也展现出巨大的潜力,有望革新药物研发、疾病诊断和治疗模式。随着技术的不断成熟和监管框架的逐步完善,虚拟细胞将在减少动物实验、优化临床试验设计、推动数字疗法等方面发挥关键作用,并创造巨大的经济和社会价值。

  1. 减少动物实验:
    传统的药物研发和毒性评估严重依赖动物实验。然而,动物实验存在伦理争议、成本高昂、耗时漫长且动物模型与人类生理差异可能导致结果外推性差等问题。虚拟细胞作为一种先进的体外(in silico)模型,提供了一个伦理上可接受、经济高效且具有更高预测准确性的替代方案。

    • 药物筛选与毒性预测: 虚拟细胞能够模拟药物在人体细胞中的作用机制、代谢过程以及潜在的毒性效应。通过在虚拟环境中筛选和评估药物的有效性和安全性,可以在进入动物实验或临床试验之前,淘汰大量无效或有毒的候选药物,显著减少所需的动物数量。例如,利用虚拟细胞模型评估药物对特定细胞类型(如肝细胞、心肌细胞)的毒性,可以有效预测药物的器官特异性毒性,从而降低药物研发的风险。
    • 疾病机制研究: 许多疾病的动物模型并不能完全复制人类疾病的复杂病理生理过程。虚拟细胞可以整合人类原代细胞或患者诱导多能干细胞(iPSC)衍生的细胞数据,构建更接近人类生理状态的疾病模型,从而减少对动物模型的依赖,加速对人类疾病机制的理解。
  2. 优化临床试验设计:
    临床试验是药物研发中最昂贵、耗时且风险最高的阶段。虚拟细胞技术有望通过提供更精准的患者分层、预测药物响应和优化试验方案,从而显著提高临床试验的效率和成功率。

    • 患者分层与选择: 虚拟细胞可以通过整合患者的基因组、蛋白质组等高维数据,构建个体化的虚拟细胞模型,预测不同患者对特定治疗的反应。这使得研究者能够更精准地筛选出最有可能从某种治疗中获益的患者群体,从而提高临床试验的成功率,降低成本。例如,在肿瘤免疫治疗中,虚拟细胞可以模拟患者肿瘤细胞与免疫细胞的相互作用,预测患者对PD-1/PD-L1抑制剂的响应,指导患者入组。
    • 虚拟患者与“合成控制组”: 虚拟细胞可以进一步扩展为“虚拟患者”模型,模拟患者群体对药物的反应。在某些情况下,这些虚拟患者数据甚至可以作为“合成控制组”,与真实治疗组进行比较,尤其在罕见病或儿科研究中,这可以减少对真实患者的招募数量,加快药物审批进程。
    • 剂量优化与方案调整: 虚拟细胞模型可以模拟不同剂量、给药方案对细胞或组织的影响,从而优化药物的剂量和给药策略,减少临床试验中探索性剂量的摸索,降低患者风险。例如,通过模拟药物在不同组织中的渗透和作用,可以优化给药途径和频率。
  3. 推动数字疗法(Digital Therapeutics, DTx):
    数字疗法是利用软件程序干预疾病管理和治疗的基于证据的治疗方案。虚拟细胞技术可以为数字疗法提供更深层次的生物学基础和个性化支持。

    • 个性化疾病管理: 虚拟细胞模型可以整合患者的实时生理数据(如通过可穿戴设备获取),模拟疾病进展和治疗响应。数字疗法可以基于这些虚拟细胞模型的预测,为患者提供个性化的健康建议、生活方式干预或治疗方案调整。例如,在糖尿病管理中,虚拟细胞模型可以预测血糖对饮食和运动的反应,数字疗法可以据此生成个性化的饮食运动计划。
    • 药物伴随诊断与优化: 虚拟细胞可以作为数字疗法的核心组件,帮助医生和患者理解药物的作用机制,预测副作用,并动态调整治疗方案。例如,在肿瘤治疗中,数字疗法结合虚拟细胞模型,可以根据患者肿瘤的动态变化,实时评估药物疗效,并提供最佳的用药建议。
  4. 政策支持与监管标准对转化的影响:
    虚拟细胞的临床转化和产业应用,离不开健全的政策支持和明确的监管标准。

    • 政策鼓励: 各国政府和研究机构已经开始认识到计算模拟在生物医学领域的潜力,并出台相关政策鼓励其发展和应用。例如,美国食品药品监督管理局(FDA)正在积极探索计算建模和模拟(CM&S)在医疗产品监管审查中的应用,并已发布指导文件,承认其在减少动物实验和优化临床试验方面的作用。
    • 监管框架的建立: 虚拟细胞模型要真正进入临床应用,必须通过严格的验证和监管审查。目前,对于虚拟细胞模型(尤其是AI驱动的模型)的验证标准、性能评估指标、可解释性要求以及数据隐私和安全性等方面的监管框架仍在发展中。建立统一的、国际认可的监管标准将是推动其广泛应用的关键。这包括定义模型的适用范围、验证方法、不确定性量化以及如何将模型结果整合到临床决策流程中。
    • 技术成熟度: 尽管虚拟细胞取得了显著进展,但其技术成熟度仍在不断提升。模型的精度、鲁棒性、可扩展性以及与真实生物系统的匹配度仍需进一步优化。尤其是在跨尺度模拟和处理高度异质性数据方面,仍有许多技术难点需要攻克。随着AI算法的进步、计算能力的提升以及生物学数据的积累,虚拟细胞的预测能力将不断增强,从而加速其临床转化。

总而言之,虚拟细胞技术正处在一个快速发展的阶段,其在减少动物实验、优化临床试验设计和推动数字疗法等方面的产业价值巨大。通过跨学科合作、开放科学实践以及政策和监管的共同推动,虚拟细胞有望成为未来医疗健康领域的重要组成部分,为人类健康福祉带来深远的影响。

内容由 AI 生成,仅供参考,请仔细甄别

参考文献

1How to build the virtual cell with artificial intelligence: Priorities and opportunities.PubMed

Charlotte Bunne, Yusuf Roohani, Yanay Rosen, et al.
Cell. 2024 Dec 12;187(25):7045-7063. doi: 10.1016/j.cell.2024.11.015.
Cells are essential to understanding health and disease, yet traditional models fall short of modeling and simulating their function and behavior. Advances in AI and omics offer groundbreaking opportunities to create an AI virtual cell (AIVC), a multi-scale, multi-modal large-neural-network-based model that can represent and simulate the behavior of molecules, cells, and tissues across diverse states. This Perspective provides a vision on their design and how collaborative efforts to build AIVCs will transform biological research by allowing high-fidelity simulations, accelerating discoveries, and guiding experimental studies, offering new opportunities for understanding cellular functions and fostering interdisciplinary collaborations in open science.

2Human interpretable grammar encodes multicellular systems biology models to democratize virtual cell laboratories.PubMed

Jeanette A I Johnson, Daniel R Bergman, Heber L Rocha, et al.
Cell. 2025 Aug 21;188(17):4711-4733.e37. doi: 10.1016/j.cell.2025.06.048. Epub 2025 Jul 26.
Cells interact as dynamically evolving ecosystems. While recent single-cell and spatial multi-omics technologies quantify individual cell characteristics, predicting their evolution requires mathematical modeling. We propose a conceptual framework-a cell behavior hypothesis grammar-that uses natural language statements (cell rules) to create mathematical models. This enables systematic integration of biological knowledge and multi-omics data to generate in silico models, enabling virtual "thought experiments" that test and expand our understanding of multicellular systems and generate new testable hypotheses. This paper motivates and describes the grammar, offers a reference implementation, and demonstrates its use in developing both de novo mechanistic models and those informed by multi-omics data. We show its potential through examples in cancer and its broader applicability in simulating brain development. This approach bridges biological, clinical, and systems biology research for mathematical modeling at scale, allowing the community to predict emergent multicellular behavior.

3[Numerical simulation of fracture healing].PubMed

Ruisen Fu, Haisheng Yang
Sheng Wu Yi Xue Gong Cheng Xue Za Zhi. 2020 Oct 25;37(5):930-935. doi: 10.7507/1001-5515.202004010.
Fracture is a common physical injury. Its healing process involves complex biological activities at tissue, cellular and molecular levels and is affected by mechanical and biological factors. Over recent years, numerical simulation methods have been widely used to explore the mechanisms of fracture healing, design fixators and develop novel treatment strategies, etc. This paper mainly recommend the numerical methods used for simulating fracture healing and their latest research progress, which helps people better understand the mechanism of fracture healing, and also provides direction and guidance for the numerical simulation research of fracture healing in the future. First, the fracture healing process and its relationship with mechanical stimulation and biological factors are described. Then, the numerical models used for simulating fracture healing (including mechano-regulatory model, biological regulatory model and mechano-biological regulatory model) and corresponding modeling techniques (mainly including agent-based techniques and fuzzy logic controlling method) were summarized in particular. Finally, the future research directions in numerical simulation of fracture healing were preliminarily prospected.

4An in vitro assay and artificial intelligence approach to determine rate constants of nanomaterial-cell interactions.PubMed

Edward Price, Andre J Gesquiere
Sci Rep. 2019 Sep 26;9(1):13943. doi: 10.1038/s41598-019-50208-x.
In vitro assays and simulation technologies are powerful methodologies that can inform scientists of nanomaterial (NM) distribution and fate in humans or pre-clinical species. For small molecules, less animal data is often needed because there are a multitude of in vitro screening tools and simulation-based approaches to quantify uptake and deliver data that makes extrapolation to in vivo studies feasible. Small molecule simulations work because these materials often diffuse quickly and partition after reaching equilibrium shortly after dosing, but this cannot be applied to NMs. NMs interact with cells through energy dependent pathways, often taking hours or days to become fully internalized within the cellular environment. In vitro screening tools must capture these phenomena so that cell simulations built on mechanism-based models can deliver relationships between exposure dose and mechanistic biology, that is biology representative of fundamental processes involved in NM transport by cells (e.g. membrane adsorption and subsequent internalization). Here, we developed, validated, and applied the FORECAST method, a combination of a calibrated fluorescence assay (CF) with an artificial intelligence-based cell simulation to quantify rates descriptive of the time-dependent mechanistic biological interactions between NMs and individual cells. This work is expected to provide a means of extrapolation to pre-clinical or human biodistribution with cellular level resolution for NMs starting only from in vitro data.

5BioUML: an integrated environment for systems biology and collaborative analysis of biomedical data.PubMed

Fedor Kolpakov, Ilya Akberdin, Timur Kashapov, et al.
Nucleic Acids Res. 2019 Jul 2;47(W1):W225-W233. doi: 10.1093/nar/gkz440.
BioUML (homepage: http://www.biouml.org, main public server: https://ict.biouml.org) is a web-based integrated environment (platform) for systems biology and the analysis of biomedical data generated by omics technologies. The BioUML vision is to provide a computational platform to build virtual cell, virtual physiological human and virtual patient. BioUML spans a comprehensive range of capabilities, including access to biological databases, powerful tools for systems biology (visual modelling, simulation, parameters fitting and analyses), a genome browser, scripting (R, JavaScript) and a workflow engine. Due to integration with the Galaxy platform and R/Bioconductor, BioUML provides powerful possibilities for the analyses of omics data. The plug-in-based architecture allows the user to add new functionalities using plug-ins. To facilitate a user focus on a particular task or database, we have developed several predefined perspectives that display only those web interface elements that are needed for a specific task. To support collaborative work on scientific projects, there is a central authentication and authorization system (https://bio-store.org). The diagram editor enables several remote users to simultaneously edit diagrams.

6All systems go: launching cell simulation fueled by integrated experimental biology data.PubMed

Masanori Arita, Martin Robert, Masaru Tomita
Curr Opin Biotechnol. 2005 Jun;16(3):344-9. doi: 10.1016/j.copbio.2005.04.004.
Biological simulation serves to unify the basic elements of systems biology, namely, model selection, experimentation and model refinement. To select biochemical models for simulation, metabolome analysis can be performed using capillary electrophoresis or liquid chromatography coupled with mass spectrometry. In this manner, selected models can be elaborated with temporal/spatial gene and protein expression data obtained from model organisms such as Escherichia coli. The E. coli single gene deletion mutant library (KO collection) and His-tag/GFP-fusion single open reading frame clone expression library (ASKA) are powerful resources for this task. The integration of parallel experimental datasets into dynamic simulation tools forms the remaining challenge for the systematic analysis and elucidation of biological networks and holds promise for biotechnological applications.

7Explaining Regeneration: Cells and Limbs as Complex Living Systems, Learning From History.PubMed

Kate MacCord, Jane Maienschein
Front Cell Dev Biol. 2021 Aug 31;9:734315. doi: 10.3389/fcell.2021.734315. eCollection 2021.
Regeneration has been investigated since Aristotle, giving rise to many ways of explaining what this process is and how it works. Current research focuses on gene expression and cell signaling of regeneration within individual model organisms. We tend to look to model organisms on the reasoning that because of evolution, information gained from other species must in some respect be generalizable. However, for all that we have uncovered about how regeneration works within individual organisms, we have yet to translate what we have gleaned into achieving the goal of regenerative medicine: to harness and enhance our own regenerative abilities. Turning to history may provide a crucial perspective in advancing us toward this goal. History gives perspective, allowing us to reflect on how our predecessors did their work and what assumptions they made, thus also revealing limitations. History, then, may show us how we can move from our current reductionist thinking focused on particular selected model organisms toward generalizations about this crucial process that operates across complex living systems and move closer to repairing our own damaged bodies.

8SGABU computational platform for multiscale modeling: Bridging the gap between education and research.PubMed

Tijana Geroski, Orestis Gkaintes, Aleksandra Vulović, et al.
Comput Methods Programs Biomed. 2024 Jan;243:107935. doi: 10.1016/j.cmpb.2023.107935. Epub 2023 Nov 22.
BACKGROUND AND OBJECTIVE: In accordance with the latest aspirations in the field of bioengineering, there is a need to create a web accessible, but powerful cloud computational platform that combines datasets and multiscale models related to bone modeling, cancer, cardiovascular diseases and tissue engineering. The SGABU platform may become a powerful information system for research and education that can integrate data, extract information, and facilitate knowledge exchange with the goal of creating and developing appropriate computing pipelines to provide accurate and comprehensive biological information from the molecular to organ level. METHODS: The datasets integrated into the platform are obtained from experimental and/or clinical studies and are mainly in tabular or image file format, including metadata. The implementation of multiscale models, is an ambitious effort of the platform to capture phenomena at different length scales, described using partial and ordinary differential equations, which are solved numerically on complex geometries with the use of the finite element method. The majority of the SGABU platform's simulation pipelines are provided as Common Workflow Language (CWL) workflows. Each of them requires creating a CWL implementation on the backend and a user-friendly interface using standard web technologies. Platform is available at https://sgabu-test.unic.kg.ac.rs/login. RESULTS: The main dashboard of the SGABU platform is divided into sections for each field of research, each one of which includes a subsection of datasets and multiscale models. The datasets can be presented in a simple form as tabular data, or using technologies such as Plotly.js for 2D plot interactivity, Kitware Paraview Glance for 3D view. Regarding the models, the usage of Docker containerization for packing the individual tools and CWL orchestration for describing inputs with validation forms and outputs with tabular views for output visualization, interactive diagrams, 3D views and animations. CONCLUSIONS: In practice, the structure of SGABU platform means that any of the integrated workflows can work equally well on any other bioengineering platform. The key advantage of the SGABU platform over similar efforts is its versatility offered with the use of modern, modular, and extensible technology for various levels of architecture.

9Prospects for Declarative Mathematical Modeling of Complex Biological Systems.PubMed

Eric Mjolsness
Bull Math Biol. 2019 Aug;81(8):3385-3420. doi: 10.1007/s11538-019-00628-7. Epub 2019 Jun 7.
Declarative modeling uses symbolic expressions to represent models. With such expressions, one can formalize high-level mathematical computations on models that would be difficult or impossible to perform directly on a lower-level simulation program, in a general-purpose programming language. Examples of such computations on models include model analysis, relatively general-purpose model reduction maps, and the initial phases of model implementation, all of which should preserve or approximate the mathematical semantics of a complex biological model. The potential advantages are particularly relevant in the case of developmental modeling, wherein complex spatial structures exhibit dynamics at molecular, cellular, and organogenic levels to relate genotype to multicellular phenotype. Multiscale modeling can benefit from both the expressive power of declarative modeling languages and the application of model reduction methods to link models across scale. Based on previous work, here we define declarative modeling of complex biological systems by defining the operator algebra semantics of an increasingly powerful series of declarative modeling languages including reaction-like dynamics of parameterized and extended objects; we define semantics-preserving implementation and semantics-approximating model reduction transformations; and we outline a "meta-hierarchy" for organizing declarative models and the mathematical methods that can fruitfully manipulate them.

10FitMultiCell: simulating and parameterizing computational models of multi-scale and multi-cellular processes.PubMed

Emad Alamoudi, Yannik Schälte, Robert Müller, et al.
Bioinformatics. 2023 Nov 1;39(11). doi: 10.1093/bioinformatics/btad674.
MOTIVATION: Biological tissues are dynamic and highly organized. Multi-scale models are helpful tools to analyse and understand the processes determining tissue dynamics. These models usually depend on parameters that need to be inferred from experimental data to achieve a quantitative understanding, to predict the response to perturbations, and to evaluate competing hypotheses. However, even advanced inference approaches such as approximate Bayesian computation (ABC) are difficult to apply due to the computational complexity of the simulation of multi-scale models. Thus, there is a need for a scalable pipeline for modeling, simulating, and parameterizing multi-scale models of multi-cellular processes. RESULTS: Here, we present FitMultiCell, a computationally efficient and user-friendly open-source pipeline that can handle the full workflow of modeling, simulating, and parameterizing for multi-scale models of multi-cellular processes. The pipeline is modular and integrates the modeling and simulation tool Morpheus and the statistical inference tool pyABC. The easy integration of high-performance infrastructure allows to scale to computationally expensive problems. The introduction of a novel standard for the formulation of parameter inference problems for multi-scale models additionally ensures reproducibility and reusability. By applying the pipeline to multiple biological problems, we demonstrate its broad applicability, which will benefit in particular image-based systems biology. AVAILABILITY AND IMPLEMENTATION: FitMultiCell is available open-source at https://gitlab.com/fitmulticell/fit.

11Analysis and numerical simulation of an inverse problem for a structured cell population dynamics model.PubMed

Frédérique Clément, Béatrice Laroche, Frédérique Robin
Math Biosci Eng. 2019 Apr 10;16(4):3018-3046. doi: 10.3934/mbe.2019150.
In this work, we study a multiscale inverse problem associated with a multi-type model for age structured cell populations. In the single type case, the model is a McKendrick-VonFoerster like equation with a mitosis-dependent death rate and potential migration at birth. In the multi-type case, the migration term results in an unidirectional motion from one type to the next, so that the boundary condition at age 0 contains an additional extrinsic contribution from the previous type. We consider the inverse problem of retrieving microscopic information (the division rates and migration proportions) from the knowledge of macroscopic information (total number of cells per layer), given the initial condition. We first show the well-posedness of the inverse problem in the single type case using a Fredholm integral equation derived from the characteristic curves, and we use a constructive approach to obtain the lattice division rate, considering either a synchronized or non-synchronized initial condition. We take advantage of the unidirectional motion to decompose the whole model into nested submodels corresponding to self-renewal equations with an additional extrinstic contribution. We again derive a Fredholm integral equation for each submodel and deduce the well-posedness of the multi-type inverse problem. In each situation, we illustrate numerically our theoretical results.

12Modelling the effectiveness of antiviral treatment strategies to prevent household transmission of acute respiratory viruses.PubMed

Hind Zaaraoui, Clarisse Schumer, Xavier Duval, et al.
PLoS Comput Biol. 2024 Dec 5;20(12):e1012573. doi: 10.1371/journal.pcbi.1012573. eCollection 2024 Dec.
Households are a major driver of transmission of acute respiratory viruses, such as SARS-CoV-2 or Influenza. Until now antiviral treatments have mostly been used as a curative treatment in symptomatic individuals. During an outbreak, more aggressive strategies involving pre- or post-exposure prophylaxis (PrEP or PEP) could be employed to further reduce the risk of severe disease but also prevent transmission to household contacts. In order to understand the effectiveness of such strategies and the factors that may modulate them, we developed a multi-scale model that follows the infection at both the individual-level (viral dynamics) and the population-level (transmission dynamics) in households. Using a simulation study we explored different antiviral treatment strategies, evaluating their effectiveness on reducing the transmission risk and the virological burden in households for a range of virus characteristics (e.g., secondary attack rate-SAR, or time to peak viral load). We found that when the index case can be identified and treated before symptom onset, both transmission and virological burden are reduced by > 75% for most SAR values and time to peak viral load, with minimal benefit to treat additionally household contacts. While treatment initiated after index symptom onset does not reduce the risk of transmission, it can still reduce the virological burden in the household, a proxy for severe disease and subsequent transmission risk outside the household. In that case optimal strategies involve treatment of both index case and household contacts as PEP, with efficacy > 50% when peak viral load occurs after symptom onset, and 30-50% otherwise. In all the considered cases, antiviral treatment strategies were optimal for SAR ranging 20-60%, and for larger household sizes. This study highlights the opportunity of antiviral drug-based interventions in households during an outbreak to minimize viral transmission and disease burden.

13Unveiling inflammatory and prehypertrophic cell populations as key contributors to knee cartilage degeneration in osteoarthritis using multi-omics data integration.PubMed

Yue Fan, Xuzhao Bian, Xiaogao Meng, et al.
Ann Rheum Dis. 2024 Jun 12;83(7):926-944. doi: 10.1136/ard-2023-224420.
OBJECTIVES: Single-cell and spatial transcriptomics analysis of human knee articular cartilage tissue to present a comprehensive transcriptome landscape and osteoarthritis (OA)-critical cell populations. METHODS: Single-cell RNA sequencing and spatially resolved transcriptomic technology have been applied to characterise the cellular heterogeneity of human knee articular cartilage which were collected from 8 OA donors, and 3 non-OA control donors, and a total of 19 samples. The novel chondrocyte population and marker genes of interest were validated by immunohistochemistry staining, quantitative real-time PCR, etc. The OA-critical cell populations were validated through integrative analyses of publicly available bulk RNA sequencing data and large-scale genome-wide association studies. RESULTS: We identified 33 cell population-specific marker genes that define 11 chondrocyte populations, including 9 known populations and 2 new populations, that is, pre-inflammatory chondrocyte population (preInfC) and inflammatory chondrocyte population (InfC). The novel findings that make this an important addition to the literature include: (1) the novel InfC activates the mediator MIF-CD74; (2) the prehypertrophic chondrocyte (preHTC) and hypertrophic chondrocyte (HTC) are potentially OA-critical cell populations; (3) most OA-associated differentially expressed genes reside in the articular surface and superficial zone; (4) the prefibrocartilage chondrocyte (preFC) population is a major contributor to the stratification of patients with OA, resulting in both an inflammatory-related subtype and a non-inflammatory-related subtype. CONCLUSIONS: Our results highlight InfC, preHTC, preFC and HTC as potential cell populations to target for therapy. Also, we conclude that profiling of those cell populations in patients might be used to stratify patient populations for defining cohorts for clinical trials and precision medicine.

14Single-Cell Multiomics Profiling Reveals Heterogeneity of Müller Cells in the Oxygen-Induced Retinopathy Model.PubMed

Xueming Yao, Ziqi Li, Yi Lei, et al.
Invest Ophthalmol Vis Sci. 2024 Nov 4;65(13):8. doi: 10.1167/iovs.65.13.8.
PURPOSE: Retinal neovascularization poses heightened risks of vision loss and blindness. Despite its clinical significance, the molecular mechanisms underlying the pathogenesis of retinal neovascularization remain elusive. This study utilized single-cell multiomics profiling in an oxygen-induced retinopathy (OIR) model to comprehensively investigate the intricate molecular landscape of retinal neovascularization. METHODS: Mice were exposed to hyperoxia to induce the OIR model, and retinas were isolated for nucleus isolation. The cellular landscape of the single-nucleus suspensions was extensively characterized through single-cell multiomics sequencing. Single-cell data were integrated with genome-wide association study (GWAS) data to identify correlations between ocular cell types and diabetic retinopathy. Cell communication analysis among cells was conducted to unravel crucial ligand-receptor signals. Trajectory analysis and dynamic characterization of Müller cells were performed, followed by integration with human retinal data for pathway analysis. RESULTS: The multiomics dataset revealed six major ocular cell classes, with Müller cells/astrocytes showing significant associations with proliferative diabetic retinopathy (PDR). Cell communication analysis highlighted pathways that are associated with vascular proliferation and neurodevelopment, such as Vegfa-Vegfr2, Igf1-Igf1r, Nrxn3-Nlgn1, and Efna5-Epha4. Trajectory analysis identified a subset of Müller cells expressing genes linked to photoreceptor degeneration. Multiomics data integration further unveiled positively regulated genes in OIR Müller cells/astrocytes associated with axon development and neurotransmitter transmission. CONCLUSIONS: This study significantly advances our understanding of the intricate cellular and molecular mechanisms underlying retinal neovascularization, emphasizing the pivotal role of Müller cells. The identified pathways provide valuable insights into potential therapeutic targets for PDR, offering promising directions for further research and clinical interventions.

15spSeudoMap: cell type mapping of spatial transcriptomics using unmatched single-cell RNA-seq data.PubMed

Sungwoo Bae, Hongyoon Choi, Dong Soo Lee
Genome Med. 2023 Mar 17;15(1):19. doi: 10.1186/s13073-023-01168-5.
Since many single-cell RNA-seq (scRNA-seq) data are obtained after cell sorting, such as when investigating immune cells, tracking cellular landscape by integrating single-cell data with spatial transcriptomic data is limited due to cell type and cell composition mismatch between the two datasets. We developed a method, spSeudoMap, which utilizes sorted scRNA-seq data to create virtual cell mixtures that closely mimic the gene expression of spatial data and trains a domain adaptation model for predicting spatial cell compositions. The method was applied in brain and breast cancer tissues and accurately predicted the topography of cell subpopulations. spSeudoMap may help clarify the roles of a few, but crucial cell types.

16iSMOD: an integrative browser for image-based single-cell multi-omics data.PubMed

Weihang Zhang, Jinli Suo, Yan Yan, et al.
Nucleic Acids Res. 2023 Sep 8;51(16):8348-8366. doi: 10.1093/nar/gkad580.
Genomic and transcriptomic image data, represented by DNA and RNA fluorescence in situ hybridization (FISH), respectively, together with proteomic data, particularly that related to nuclear proteins, can help elucidate gene regulation in relation to the spatial positions of chromatins, messenger RNAs, and key proteins. However, methods for image-based multi-omics data collection and analysis are lacking. To this end, we aimed to develop the first integrative browser called iSMOD (image-based Single-cell Multi-omics Database) to collect and browse comprehensive FISH and nucleus proteomics data based on the title, abstract, and related experimental figures, which integrates multi-omics studies focusing on the key players in the cell nucleus from 20 000+ (still growing) published papers. We have also provided several exemplar demonstrations to show iSMOD's wide applications-profiling multi-omics research to reveal the molecular target for diseases; exploring the working mechanism behind biological phenomena using multi-omics interactions, and integrating the 3D multi-omics data in a virtual cell nucleus. iSMOD is a cornerstone for delineating a global view of relevant research to enable the integration of scattered data and thus provides new insights regarding the missing components of molecular pathway mechanisms and facilitates improved and efficient scientific research.

17The Quality Assurance and Quality Control Protocol for Neuropsychological Data Collection and Curation in the Ontario Neurodegenerative Disease Research Initiative (ONDRI) Study.PubMed

Paula M McLaughlin, Kelly M Sunderland, Derek Beaton, et al.
Assessment. 2021 Jul;28(5):1267-1286. doi: 10.1177/1073191120913933. Epub 2020 Apr 22.
As large research initiatives designed to generate big data on clinical cohorts become more common, there is an increasing need to establish standard quality assurance (QA; preventing errors) and quality control (QC; identifying and correcting errors) procedures for critical outcome measures. The present article describes the QA and QC approach developed and implemented for the neuropsychology data collected as part of the Ontario Neurodegenerative Disease Research Initiative study. We report on the efficacy of our approach and provide data quality metrics. Our findings demonstrate that even with a comprehensive QA protocol, the proportion of data errors still can be high. Additionally, we show that several widely used neuropsychological measures are particularly susceptible to error. These findings highlight the need for large research programs to put into place active, comprehensive, and separate QA and QC procedures before, during, and after protocol deployment. Detailed recommendations and considerations for future studies are provided.

18How to Build the Virtual Cell with Artificial Intelligence: Priorities and Opportunities.PubMed

Charlotte Bunne, Yusuf Roohani, Yanay Rosen, et al.
ArXiv. 2024 Oct 14:arXiv:2409.11654v2.
The cell is arguably the most fundamental unit of life and is central to understanding biology. Accurate modeling of cells is important for this understanding as well as for determining the root causes of disease. Recent advances in artificial intelligence (AI), combined with the ability to generate large-scale experimental data, present novel opportunities to model cells. Here we propose a vision of leveraging advances in AI to construct virtual cells, high-fidelity simulations of cells and cellular systems under different conditions that are directly learned from biological data across measurements and scales. We discuss desired capabilities of such AI Virtual Cells, including generating universal representations of biological entities across scales, and facilitating interpretable experiments to predict and understand their behavior using Virtual Instruments. We further address the challenges, opportunities and requirements to realize this vision including data needs, evaluation strategies, and community standards and engagement to ensure biological accuracy and broad utility. We envision a future where AI Virtual Cells help identify new drug targets, predict cellular responses to perturbations, as well as scale hypothesis exploration. With open science collaborations across the biomedical ecosystem that includes academia, philanthropy, and the biopharma and AI industries, a comprehensive predictive understanding of cell mechanisms and interactions has come into reach.

19The cell as a token: high-dimensional geometry in language models and cell embeddings.PubMed

William Gilpin
Bioinformatics. 2025 Nov 1;41(11). doi: 10.1093/bioinformatics/btaf595.
MOTIVATION: Single-cell sequencing technology maps cells to a high-dimensional space encoding their internal activity. Recently-proposed virtual cell models extend this concept, enriching cells' representations based on patterns learned from pretraining on vast cell atlases. RESULTS: This review explores how advances in understanding the structure of natural language embeddings informs ongoing efforts to analyze single-cell datasets. Both fields process unstructured data by partitioning datasets into tokens embedded within a high-dimensional vector space. We discuss how the context of tokens influences the geometry of embedding space, and how low-dimensional manifolds shape this space's robustness and interpretation. We highlight how new developments in foundation models for language, such as interpretability probes and in-context reasoning, can inform efforts to construct cell atlases and train virtual cell models. AVAILABILITY AND IMPLEMENTATION: Code is available at https://github.com/williamgilpin/celltoken.

20Applying Spatiotemporal Modeling of Cell Dynamics to Accelerate Drug Development.PubMed

Xindong Chen, Shihao Xu, Bizhu Chu, et al.
ACS Nano. 2024 Oct 29;18(43):29311-29336. doi: 10.1021/acsnano.4c12599. Epub 2024 Oct 18.
Cells act as physical computational programs that utilize input signals to orchestrate molecule-level protein-protein interactions (PPIs), generating and responding to forces, ultimately shaping all of the physiological and pathophysiological behaviors. Genome editing and molecule drugs targeting PPIs hold great promise for the treatments of diseases. Linking genes and molecular drugs with protein-performed cellular behaviors is a key yet challenging issue due to the wide range of spatial and temporal scales involved. Building predictive spatiotemporal modeling systems that can describe the dynamic behaviors of cells intervened by genome editing and molecular drugs at the intersection of biology, chemistry, physics, and computer science will greatly accelerate pharmaceutical advances. Here, we review the mechanical roles of cytoskeletal proteins in orchestrating cellular behaviors alongside significant advancements in biophysical modeling while also addressing the limitations in these models. Then, by integrating generative artificial intelligence (AI) with spatiotemporal multiscale biophysical modeling, we propose a computational pipeline for developing virtual cells, which can simulate and evaluate the therapeutic effects of drugs and genome editing technologies on various cell dynamic behaviors and could have broad biomedical applications. Such virtual cell modeling systems might revolutionize modern biomedical engineering by moving most of the painstaking wet-laboratory effort to computer simulations, substantially saving time and alleviating the financial burden for pharmaceutical industries.

21AI-driven virtual cell models in preclinical research: technical pathways, validation mechanisms, and clinical translation potential.PubMed

Chunyu Ma, Han Zhang, Yiwei Rao, et al.
NPJ Digit Med. 2025 Dec 11;9(1):25. doi: 10.1038/s41746-025-02198-6.
AI-driven virtual cell models show the potential to transform the paradigm of life sciences research by integrating multimodal omics data (e.g., single-cell transcriptomics and proteomics) with advanced algorithms such as deep generative models and graph neural networks to enable high-precision predictions of drug responses, gene perturbations, and disease progression. These models enable high-precision predictions of drug responses, gene perturbations, and disease progression. This review outlines the technical pathways and validation mechanisms of virtual cells, emphasizing a closed-loop workflow from computational evaluation to experimental verification using CRISPR assays and organoid platforms. The applications of virtual cells in personalized drug screening and disease modeling are highlighted, showcasing their potential to reduce animal testing and optimize therapy. However, challenges in regulatory acceptance, data privacy, and model interpretability remain. Global policy and standardization trends are driving clinical translation, and future advancements will involve cross-disciplinary integration and greater standardization to enhance the impact of virtual cells in precision medicine and drug discovery.

22Empowering biomedical discovery with AI agents.PubMed

Shanghua Gao, Ada Fang, Yepeng Huang, et al.
Cell. 2024 Oct 31;187(22):6125-6151. doi: 10.1016/j.cell.2024.09.022.
We envision "AI scientists" as systems capable of skeptical learning and reasoning that empower biomedical research through collaborative agents that integrate AI models and biomedical tools with experimental platforms. Rather than taking humans out of the discovery process, biomedical AI agents combine human creativity and expertise with AI's ability to analyze large datasets, navigate hypothesis spaces, and execute repetitive tasks. AI agents are poised to be proficient in various tasks, planning discovery workflows and performing self-assessment to identify and mitigate gaps in their knowledge. These agents use large language models and generative models to feature structured memory for continual learning and use machine learning tools to incorporate scientific knowledge, biological principles, and theories. AI agents can impact areas ranging from virtual cell simulation, programmable control of phenotypes, and the design of cellular circuits to developing new therapies.

23Invited Review for 20th Anniversary Special Issue of PLRev "AI for Mechanomedicine".PubMed

Ning Xie, Jin Tian, Zedong Li, et al.
Phys Life Rev. 2024 Dec;51:328-342. doi: 10.1016/j.plrev.2024.10.010. Epub 2024 Oct 24.
Mechanomedicine is an interdisciplinary field that combines different areas including biomechanics, mechanobiology, and clinical applications like mechanodiagnosis and mechanotherapy. The emergence of artificial intelligence (AI) has revolutionized mechanomedicine, providing advanced tools to analyze the complex interactions between mechanics and biology. This review explores how AI impacts mechanomedicine across four key aspects, i.e., biomechanics, mechanobiology, mechanodiagnosis, and mechanotherapy. AI improves the accuracy of biomechanical characterizations and models, deepens the understanding of cellular mechanotransduction pathways, and enables early disease detection through mechanodiagnosis. In addition, AI optimizes mechanotherapy that targets biomechanical features and mechanobiological markers by personalizing treatment strategies based on real-time patient data. Even with these advancements, challenges still exist, particularly in data quality and the ethical integration into AI in clinical practice. The integration of AI with mechanomedicine offers transformative potential, enabling more accurate diagnostics and personalized treatments, and discovering novel mechanobiological pathways.

24Spatial modeling of cell signaling networks.PubMed

Ann E Cowan, Ion I Moraru, James C Schaff, et al.
Methods Cell Biol. 2012;110:195-221. doi: 10.1016/B978-0-12-388403-9.00008-4.
The shape of a cell, the sizes of subcellular compartments, and the spatial distribution of molecules within the cytoplasm can all control how molecules interact to produce a cellular behavior. This chapter describes how these spatial features can be included in mechanistic mathematical models of cell signaling. The Virtual Cell computational modeling and simulation software is used to illustrate the considerations required to build a spatial model. An explanation of how to appropriately choose between physical formulations that implicitly or explicitly account for cell geometry and between deterministic versus stochastic formulations for molecular dynamics is provided, along with a discussion of their respective strengths and weaknesses. As a first step toward constructing a spatial model, the geometry needs to be specified and associated with the molecules, reactions, and membrane flux processes of the network. Initial conditions, diffusion coefficients, velocities, and boundary conditions complete the specifications required to define the mathematics of the model. The numerical methods used to solve reaction-diffusion problems both deterministically and stochastically are then described and some guidance is provided in how to set up and run simulations. A study of cAMP signaling in neurons ends the chapter, providing an example of the insights that can be gained in interpreting experimental results through the application of spatial modeling.

25Rule-based modeling with Virtual Cell.PubMed

James C Schaff, Dan Vasilescu, Ion I Moraru, et al.
Bioinformatics. 2016 Sep 15;32(18):2880-2. doi: 10.1093/bioinformatics/btw353. Epub 2016 Jun 9.
UNLABELLED: Rule-based modeling is invaluable when the number of possible species and reactions in a model become too large to allow convenient manual specification. The popular rule-based software tools BioNetGen and NFSim provide powerful modeling and simulation capabilities at the cost of learning a complex scripting language which is used to specify these models. Here, we introduce a modeling tool that combines new graphical rule-based model specification with existing simulation engines in a seamless way within the familiar Virtual Cell (VCell) modeling environment. A mathematical model can be built integrating explicit reaction networks with reaction rules. In addition to offering a large choice of ODE and stochastic solvers, a model can be simulated using a network free approach through the NFSim simulation engine. AVAILABILITY AND IMPLEMENTATION: Available as VCell (versions 6.0 and later) at the Virtual Cell web site (http://vcell.org/). The application installs and runs on all major platforms and does not require registration for use on the user's computer. Tutorials are available at the Virtual Cell website and Help is provided within the software. Source code is available at Sourceforge. CONTACT: vcell_support@uchc.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

26Pathway Commons at virtual cell: use of pathway data for mathematical modeling.PubMed

Michael L Blinov, James C Schaff, Oliver Ruebenacker, et al.
Bioinformatics. 2014 Jan 15;30(2):292-4. doi: 10.1093/bioinformatics/btt660. Epub 2013 Nov 22.
UNLABELLED: Pathway Commons is a resource permitting simultaneous queries of multiple pathway databases. However, there is no standard mechanism for using these data (stored in BioPAX format) to annotate and build quantitative mathematical models. Therefore, we developed a new module within the virtual cell modeling and simulation software. It provides pathway data retrieval and visualization and enables automatic creation of executable network models directly from qualitative connections between pathway nodes. AVAILABILITY AND IMPLEMENTATION: Available at Virtual Cell (http://vcell.org/). Application runs on all major platforms and does not require registration for use on the user’s computer. Tutorials and video are available at user guide page.

27Modeling gene expression networks using fuzzy logic.PubMed

Pan Du, Jian Gong, Eve Syrkin Wurtele, et al.
IEEE Trans Syst Man Cybern B Cybern. 2005 Dec;35(6):1351-9. doi: 10.1109/tsmcb.2005.855590.
Gene regulatory networks model regulation in living organisms. Fuzzy logic can effectively model gene regulation and interaction to accurately reflect the underlying biology. A new multiscale fuzzy clustering method allows genes to interact between regulatory pathways and across different conditions at different levels of detail. Fuzzy cluster centers can be used to quickly discover causal relationships between groups of coregulated genes. Fuzzy measures weight expert knowledge and help quantify uncertainty about the functions of genes using annotations and the gene ontology database to confirm some of the interactions. The method is illustrated using gene expression data from an experiment on carbohydrate metabolism in the model plant Arabidopsis thaliana. Key gene regulatory relationships were evaluated using information from the gene ontology database. A new regulatory relationship concerning trehalose regulation of carbohydrate metabolism was also discovered in the extracted network.

28Reconstruction of biological networks based on life science data integration.PubMed

Benjamin Kormeier, Klaus Hippe, Patrizio Arrigo, et al.
J Integr Bioinform. 2010 Oct 27;7(2):428. doi: 10.2390/biecoll-jib-2010-146.
For the implementation of the virtual cell, the fundamental question is how to model and simulate complex biological networks. Therefore, based on relevant molecular database and information systems, biological data integration is an essential step in constructing biological networks. In this paper, we will motivate the applications BioDWH--an integration toolkit for building life science data warehouses, CardioVINEdb--a information system for biological data in cardiovascular-disease and VANESA--a network editor for modeling and simulation of biological networks. Based on this integration process, the system supports the generation of biological network models. A case study of a cardiovascular-disease related gene-regulated biological network is also presented.

29Internet of Bio Nano Things-based FRET nanocommunications for eHealth.PubMed

Saied M Abd El-Atty, Konstantinos A Lizos, Osama Alfarraj, et al.
Math Biosci Eng. 2023 Mar 15;20(5):9246-9267. doi: 10.3934/mbe.2023405.
The integration of the Internet of Bio Nano Things (IoBNT) with artificial intelligence (AI) and molecular communications technology is now required to achieve eHealth, specifically in the targeted drug delivery system (TDDS). In this work, we investigate an analytical framework for IoBNT with Forster resonance energy transfer (FRET) nanocommunication to enable intelligent bio nano thing (BNT) machine to accurately deliver therapeutic drug to the diseased cells. The FRET nanocommunication is accomplished by using the well-known pair of fluorescent proteins, EYFP and ECFP. Furthermore, the proposed IoBNT monitors drug transmission by using the quenching process in order to reduce side effects in healthy cells. We investigate the IoBNT framework by driving diffusional rate models in the presence of a quenching process. We evaluate the performance of the proposed framework in terms of the energy transfer efficiency, diffusion-controlled rate and drug loss rate. According to the simulation results, the proposed IoBNT with the intelligent bio nano thing for monitoring the quenching process can significantly achieve high energy transfer efficiency and low drug delivery loss rate, i.e., accurately delivering the desired therapeutic drugs to the diseased cell.

30Verification, validation and sensitivity studies in computational biomechanics.PubMed

Andrew E Anderson, Benjamin J Ellis, Jeffrey A Weiss
Comput Methods Biomech Biomed Engin. 2007 Jun;10(3):171-84. doi: 10.1080/10255840601160484.
Computational techniques and software for the analysis of problems in mechanics have naturally moved from their origins in the traditional engineering disciplines to the study of cell, tissue and organ biomechanics. Increasingly complex models have been developed to describe and predict the mechanical behavior of such biological systems. While the availability of advanced computational tools has led to exciting research advances in the field, the utility of these models is often the subject of criticism due to inadequate model verification and validation (V&V). The objective of this review is to present the concepts of verification, validation and sensitivity studies with regard to the construction, analysis and interpretation of models in computational biomechanics. Specific examples from the field are discussed. It is hoped that this review will serve as a guide to the use of V&V principles in the field of computational biomechanics, thereby improving the peer acceptance of studies that use computational modeling techniques.

31Chemical genomics with pyrvinium identifies C1orf115 as a regulator of drug efflux.PubMed

Sanna N Masud, Megha Chandrashekhar, Michael Aregger, et al.
Nat Chem Biol. 2022 Dec;18(12):1370-1379. doi: 10.1038/s41589-022-01109-0. Epub 2022 Aug 15.
Pyrvinium is a quinoline-derived cyanine dye and an approved anti-helminthic drug reported to inhibit WNT signaling and have anti-proliferative effects in various cancer cell lines. To further understand the mechanism by which pyrvinium is cytotoxic, we conducted a pooled genome-wide CRISPR loss-of-function screen in the human HAP1 cell model. The top drug-gene sensitizer interactions implicated the malate-aspartate and glycerol-3-phosphate shuttles as mediators of cytotoxicity to mitochondrial complex I inhibition including pyrvinium. By contrast, perturbation of the poorly characterized gene C1orf115/RDD1 resulted in strong resistance to the cytotoxic effects of pyrvinium through dysregulation of the major drug efflux pump ABCB1/MDR1. Interestingly, C1orf115/RDD1 was found to physically associate with ABCB1/MDR1 through proximity-labeling experiments and perturbation of C1orf115 led to mis-localization of ABCB1/MDR1. Our results are consistent with a model whereby C1orf115 modulates drug efflux through regulation of the major drug exporter ABCB1/MDR1.

32Pooled Lentiviral-Delivery Genetic Screens.PubMed

Federica Piccioni, Scott T Younger, David E Root
Curr Protoc Mol Biol. 2018 Jan 16;121:32.1.1-32.1.21. doi: 10.1002/cpmb.52.
Pooled cell-based screens of mammalian genetic perturbations enable systematic large-scale, even genome-scale, evaluation of gene function. Pooled screens introduce genetic perturbations into a cell population through viral transduction such that each cell integrates into its DNA a single or small number of library perturbations with barcodes identifying the perturbations. One then selects and physically isolates the subset of cells that exhibit the phenotype of interest. Sequencing the barcodes in the hit cells reveals which genes favored or inhibited the hit phenotype. Various genetic perturbations are possible, including CRISPR gene knockout, ectopic gene expression, and RNA interference. Regardless of the type of library being screened or the type of cell model being tested, such screens involve many common steps and procedures. This unit describes detailed experimental protocols for the key steps, and also highlights some of the key factors to achieving a well-powered, reproducible screen result. © 2018 by John Wiley & Sons, Inc.

33Human Bone Marrow Organoids for Disease Modeling, Discovery, and Validation of Therapeutic Targets in Hematologic Malignancies.PubMed

Abdullah O Khan, Antonio Rodriguez-Romera, Jasmeet S Reyat, et al.
Cancer Discov. 2023 Feb 6;13(2):364-385. doi: 10.1158/2159-8290.CD-22-0199.
UNLABELLED: A lack of models that recapitulate the complexity of human bone marrow has hampered mechanistic studies of normal and malignant hematopoiesis and the validation of novel therapies. Here, we describe a step-wise, directed-differentiation protocol in which organoids are generated from induced pluripotent stem cells committed to mesenchymal, endothelial, and hematopoietic lineages. These 3D structures capture key features of human bone marrow-stroma, lumen-forming sinusoids, and myeloid cells including proplatelet-forming megakaryocytes. The organoids supported the engraftment and survival of cells from patients with blood malignancies, including cancer types notoriously difficult to maintain ex vivo. Fibrosis of the organoid occurred following TGFβ stimulation and engraftment with myelofibrosis but not healthy donor-derived cells, validating this platform as a powerful tool for studies of malignant cells and their interactions within a human bone marrow-like milieu. This enabling technology is likely to accelerate the discovery and prioritization of novel targets for bone marrow disorders and blood cancers. SIGNIFICANCE: We present a human bone marrow organoid that supports the growth of primary cells from patients with myeloid and lymphoid blood cancers. This model allows for mechanistic studies of blood cancers in the context of their microenvironment and provides a much-needed ex vivo tool for the prioritization of new therapeutics. See related commentary by Derecka and Crispino, p. 263. This article is highlighted in the In This Issue feature, p. 247.

34A human breast cancer-derived xenograft and organoid platform for drug discovery and precision oncology.PubMed

Katrin P Guillen, Maihi Fujita, Andrew J Butterfield, et al.
Nat Cancer. 2022 Feb;3(2):232-250. doi: 10.1038/s43018-022-00337-6. Epub 2022 Feb 24.
Models that recapitulate the complexity of human tumors are urgently needed to develop more effective cancer therapies. We report a bank of human patient-derived xenografts (PDXs) and matched organoid cultures from tumors that represent the greatest unmet need: endocrine-resistant, treatment-refractory and metastatic breast cancers. We leverage matched PDXs and PDX-derived organoids (PDxO) for drug screening that is feasible and cost-effective with in vivo validation. Moreover, we demonstrate the feasibility of using these models for precision oncology in real time with clinical care in a case of triple-negative breast cancer (TNBC) with early metastatic recurrence. Our results uncovered a Food and Drug Administration (FDA)-approved drug with high efficacy against the models. Treatment with this therapy resulted in a complete response for the individual and a progression-free survival (PFS) period more than three times longer than their previous therapies. This work provides valuable methods and resources for functional precision medicine and drug development for human breast cancer.

35Naturalistic neuroscience and virtual reality.PubMed

Kay Thurley
Front Syst Neurosci. 2022 Nov 17;16:896251. doi: 10.3389/fnsys.2022.896251. eCollection 2022.
Virtual reality (VR) is one of the techniques that became particularly popular in neuroscience over the past few decades. VR experiments feature a closed-loop between sensory stimulation and behavior. Participants interact with the stimuli and not just passively perceive them. Several senses can be stimulated at once, large-scale environments can be simulated as well as social interactions. All of this makes VR experiences more natural than those in traditional lab paradigms. Compared to the situation in field research, a VR simulation is highly controllable and reproducible, as required of a laboratory technique used in the search for neural correlates of perception and behavior. VR is therefore considered a middle ground between ecological validity and experimental control. In this review, I explore the potential of VR in eliciting naturalistic perception and behavior in humans and non-human animals. In this context, I give an overview of recent virtual reality approaches used in neuroscientific research.

36Neuromechanical simulation.PubMed

Donald H Edwards
Front Behav Neurosci. 2010 Jul 14;4. doi: 10.3389/fnbeh.2010.00040. eCollection 2010.
The importance of the interaction between the body and the brain for the control of behavior has been recognized in recent years with the advent of neuromechanics, a field in which the coupling between neural and biomechanical processes is an explicit focus. A major tool used in neuromechanics is simulation, which connects computational models of neural circuits to models of an animal's body situated in a virtual physical world. This connection closes the feedback loop that links the brain, the body, and the world through sensory stimuli, muscle contractions, and body movement. Neuromechanical simulations enable investigators to explore the dynamical relationships between the brain, the body, and the world in ways that are difficult or impossible through experiment alone. Studies in a variety of animals have permitted the analysis of extremely complex and dynamic neuromechanical systems, they have demonstrated that the nervous system functions synergistically with the mechanical properties of the body, they have examined hypotheses that are difficult to test experimentally, and they have explored the role of sensory feedback in controlling complex mechanical systems with many degrees of freedom. Each of these studies confronts a common set of questions: (i) how to abstract key features of the body, the world and the CNS in a useful model, (ii) how to ground model parameters in experimental reality, (iii) how to optimize the model and identify points of sensitivity and insensitivity, and (iv) how to share neuromechanical models for examination, testing, and extension by others.

37Improved hierarchical parameter optimization technique: application for a cardiac myocyte model.PubMed

Yukiko Yamashita, Koji Sakai, Naohisa Sakamoto, et al.
Conf Proc IEEE Eng Med Biol Soc. 2006;2006:3487-90. doi: 10.1109/IEMBS.2006.260384.
We propose a hierarchical parameter optimization technique which further enhances our original technique for an accurate biological cell simulation. Our original technique generates a k-d tree and uses the coefficient of multiple determination (R2) for tree-branching, however it requires a huge computation time since it processes the response surface at every leaf node of the k-d tree, including those that do not have optimum parameters. In our parameter optimization problem, the objective function is defined as the difference between the measured and calculated waveform of action potentials in a cardiac myocyte. The function value is always non-negative, and is equal to zero if and only if the best optimized parameter is included in the leaf node. To minimize the computational cost problem, our proposed technique takes advantage of the aforementioned conditions and only processes a leaf node if the corresponding Hessian matrix of the objective function is found to be a positive definite matrix. We confirmed the effectiveness of the proposed parameter optimization technique by searching for some pre-determined parameters.

38Closing the loop: Teaching single-cell foundation models to learn from perturbations.PubMed

Yash Pershad, Tarak N Nandi, Joseph C Van Amburg, et al.
bioRxiv. 2025 Jul 12:2025.07.08.663754. doi: 10.1101/2025.07.08.663754.
The application of transfer learning models to large scale single-cell datasets has enabled the development of single-cell foundation models (scFMs) that can predict cellular responses to perturbations in silico. Although these predictions can be experimentally tested, current scFMs are unable to "close the loop" and learn from these experiments to create better predictions. Here, we introduce a "closed-loop" framework that extends the scFM by incorporating perturbation data during model fine-tuning. Our closed-loop model improves prediction accuracy, increasing positive predictive value in the setting of T-cell activation three-fold. We applied this model to RUNX1-familial platelet disorder, a rare pediatric blood disorder and identified two therapeutic targets (mTOR and CD74-MIF signaling axis) and two novel pathways (protein kinase C and phosphoinositide 3-kinase). This work establishes that iterative incorporation of experimental data to foundation models enhances biological predictions, representing a crucial step toward realizing the promise of "virtual cell" models for biomedical discovery.

39Single-cell profiling of SLC family transporters: uncovering the role of SLC7A1 in osteosarcoma.PubMed

Yan Liao, Junkai Chen, Hao Yao, et al.
J Transl Med. 2025 Jan 22;23(1):103. doi: 10.1186/s12967-025-06086-1.
BACKGROUND: Osteosarcoma is the most common malignant bone tumor in children and adolescents, characterized by high disability and mortality rates. Over the past three decades, therapeutic outcomes have plateaued, underscoring the critical need for innovative therapeutic targets. Solute carrier (SLC) family transporters have been implicated in the malignant progression of a variety of tumors, however, their specific role in osteosarcoma remains poorly understood. METHODS: The single-cell sequencing data from GSE152048 and GSE162454, along with RNA-seq from the TARGET and GSE21257 cohorts, were utilized for the analysis in this study. LASSO regression analysis was conducted to identify prognostic genes and construct an SLC-related prognostic signature. Survival analysis and ROC analysis evaluated the validity of the prognostic signature. The ESTIMATE and CIBERSORT Packages were utilized to assess the immune infiltration status. Pseudotime and CellChat analyses were performed to investigate the relationship between SLC7A1, malignant phenotypes, and the immune microenvironment. CCK8 assays, EdU staining, colony formation assays, Transwell assays, and co-culture systems were used to assess the effects of SLC7A1 on cell proliferation, metastasis, and macrophage polarization. Finally, virtual docking identified potential drugs targeting SLC7A1. RESULTS: SLCs displayed distinct expression patterns across various cell types within the osteosarcoma microenvironment, with myeloid cells exhibiting a preference for amino acid uptake. A prognostic model comprising nine genes was constructed via LASSO regression, with SLC7A1 showing the highest hazard ratio. Multiple analytical algorithms indicated that SLCs were associated with immune cell infiltration and immune checkpoint gene expression. Single-cell analysis indicated that SLC7A1 was predominantly expressed in osteosarcoma cells and correlated with various malignant tumor characteristics. SLC7A1 also regulate interactions between tumor cells and macrophages, as well as modulate macrophage function through multiple pathways. In vitro assays and survival analysis demonstrated that inhibition of SLC7A1 suppressed the malignant phenotype of osteosarcoma cells, with SLC7A1 expression correlating with poor prognosis. Co-culture models confirmed the involvement of SLC7A1 in macrophage polarization. Finally, virtual screening and CETSA identified Cepharanthine as potential inhibitors of SLC7A1. CONCLUSION: SLC-related prognostic signatures can be utilized for the prognostic evaluation of osteosarcoma. Pharmacological inhibition of SLC7A1 may be a feasible therapeutic approach for osteosarcoma.

40Prioritization of cancer therapeutic targets using CRISPR-Cas9 screens.PubMed

Fiona M Behan, Francesco Iorio, Gabriele Picco, et al.
Nature. 2019 Apr;568(7753):511-516. doi: 10.1038/s41586-019-1103-9. Epub 2019 Apr 10.
Functional genomics approaches can overcome limitations-such as the lack of identification of robust targets and poor clinical efficacy-that hamper cancer drug development. Here we performed genome-scale CRISPR-Cas9 screens in 324 human cancer cell lines from 30 cancer types and developed a data-driven framework to prioritize candidates for cancer therapeutics. We integrated cell fitness effects with genomic biomarkers and target tractability for drug development to systematically prioritize new targets in defined tissues and genotypes. We verified one of our most promising dependencies, the Werner syndrome ATP-dependent helicase, as a synthetic lethal target in tumours from multiple cancer types with microsatellite instability. Our analysis provides a resource of cancer dependencies, generates a framework to prioritize cancer drug targets and suggests specific new targets. The principles described in this study can inform the initial stages of drug development by contributing to a new, diverse and more effective portfolio of cancer drug targets.

41A comprehensive clinically informed map of dependencies in cancer cells and framework for target prioritization.PubMed

Clare Pacini, Emma Duncan, Emanuel Gonçalves, et al.
Cancer Cell. 2024 Feb 12;42(2):301-316.e9. doi: 10.1016/j.ccell.2023.12.016. Epub 2024 Jan 11.
Genetic screens in cancer cell lines inform gene function and drug discovery. More comprehensive screen datasets with multi-omics data are needed to enhance opportunities to functionally map genetic vulnerabilities. Here, we construct a second-generation map of cancer dependencies by annotating 930 cancer cell lines with multi-omic data and analyze relationships between molecular markers and cancer dependencies derived from CRISPR-Cas9 screens. We identify dependency-associated gene expression markers beyond driver genes, and observe many gene addiction relationships driven by gain of function rather than synthetic lethal effects. By combining clinically informed dependency-marker associations with protein-protein interaction networks, we identify 370 anti-cancer priority targets for 27 cancer types, many of which have network-based evidence of a functional link with a marker in a cancer type. Mapping these targets to sequenced tumor cohorts identifies tractable targets in different cancer types. This target prioritization map enhances understanding of gene dependencies and identifies candidate anti-cancer targets for drug development.

42Predicting transcriptional responses to novel chemical perturbations using deep generative model for drug discovery.PubMed

Xiaoning Qi, Lianhe Zhao, Chenyu Tian, et al.
Nat Commun. 2024 Oct 26;15(1):9256. doi: 10.1038/s41467-024-53457-1.
Understanding transcriptional responses to chemical perturbations is central to drug discovery, but exhaustive experimental screening of disease-compound combinations is unfeasible. To overcome this limitation, here we introduce PRnet, a perturbation-conditioned deep generative model that predicts transcriptional responses to novel chemical perturbations that have never experimentally perturbed at bulk and single-cell levels. Evaluations indicate that PRnet outperforms alternative methods in predicting responses across novel compounds, pathways, and cell lines. PRnet enables gene-level response interpretation and in-silico drug screening for diseases based on gene signatures. PRnet further identifies and experimentally validates novel compound candidates against small cell lung cancer and colorectal cancer. Lastly, PRnet generates a large-scale integration atlas of perturbation profiles, covering 88 cell lines, 52 tissues, and various compound libraries. PRnet provides a robust and scalable candidate recommendation workflow and successfully recommends drug candidates for 233 diseases. Overall, PRnet is an effective and valuable tool for gene-based therapeutics screening.

43AI-powered omics-based drug pair discovery for pyroptosis therapy targeting triple-negative breast cancer.PubMed

Boshu Ouyang, Caihua Shan, Shun Shen, et al.
Nat Commun. 2024 Aug 30;15(1):7560. doi: 10.1038/s41467-024-51980-9.
Due to low success rates and long cycles of traditional drug development, the clinical tendency is to apply omics techniques to reveal patient-level disease characteristics and individualized responses to treatment. However, the heterogeneous form of data and uneven distribution of targets make drug discovery and precision medicine a non-trivial task. This study takes pyroptosis therapy for triple-negative breast cancer (TNBC) as a paradigm and uses data mining of a large TNBC cohort and drug databases to establish a biofactor-regulated neural network for rapidly screening and optimizing compound pyroptosis drug pairs. Subsequently, biomimetic nanococrystals are prepared using the preferred combination of mitoxantrone and gambogic acid for rational drug delivery. The unique mechanism of obtained nanococrystals regulating pyroptosis genes through ribosomal stress and triggering pyroptosis cascade immune effects are revealed in TNBC models. In this work, a target omics-based intelligent compound drug discovery framework explores an innovative drug development paradigm, which repurposes existing drugs and enables precise treatment of refractory diseases.

44Automated workflows for modelling chemical fate, kinetics and toxicity.PubMed

J V Sala Benito, Alicia Paini, Andrea-Nicole Richarz, et al.
Toxicol In Vitro. 2017 Dec;45(Pt 2):249-257. doi: 10.1016/j.tiv.2017.03.004. Epub 2017 Mar 18.
Automation is universal in today's society, from operating equipment such as machinery, in factory processes, to self-parking automobile systems. While these examples show the efficiency and effectiveness of automated mechanical processes, automated procedures that support the chemical risk assessment process are still in their infancy. Future human safety assessments will rely increasingly on the use of automated models, such as physiologically based kinetic (PBK) and dynamic models and the virtual cell based assay (VCBA). These biologically-based models will be coupled with chemistry-based prediction models that also automate the generation of key input parameters such as physicochemical properties. The development of automated software tools is an important step in harmonising and expediting the chemical safety assessment process. In this study, we illustrate how the KNIME Analytics Platform can be used to provide a user-friendly graphical interface for these biokinetic models, such as PBK models and VCBA, which simulates the fate of chemicals in vivo within the body and in vitro test systems respectively.

45In silico reconstructions underpin aberrant trafficking dynamics in deficient axons of Dst knockout and Dst/Nefl double-knockout mice.PubMed

Zongmin Liu, Wei Wang, Elena Zhang, et al.
Commun Biol. 2025 Sep 24;8(1):1358. doi: 10.1038/s42003-025-08728-y.
Aberrant neuronal trafficking is a significant hallmark of neurodegenerative pathology. Its real-time evolution remains elusive and poorly defined due to the lack of a predictive spatiotemporal framework. Building upon a general neurocytoskeletal-PDEs (iGCPs) model, we propose the concept of Virtual Cellular Dynamics for quantitative spatiotemporal simulations of mitochondrial dynamics within axons. The model integrates interactions of key cytoskeletal components such as dystonin, microtubule, neurofilament, and actin filament, providing a comprehensive framework for neuron-specific virtual cell modeling, enabling quantitative insight into axonal dysfunction and structural degradation across neurodegenerative disease. Not only does our model recapitulate the significant structural deformations and mitochondrial transport disruptions observed in Dst-deficient mice, but it further predicts that the ablation of Nefl alleviates severe neurodegenerative progression-a finding substantiated by multi-modal imaging and Dst/Nefl double-knockout murine models, which reveal phenotypic rescue and validate the potential of NF-L-targeted therapeutic strategies. Altogether, our work paves the way for next-generation virtual cell models tailored to neuron-specific disease states.

46Integration of Immune Cell-Target Cell Conjugate Dynamics Changes the Time Scale of Immune Control of Cancer.PubMed

Qianci Yang, Arne Traulsen, Philipp M Altrock
Bull Math Biol. 2025 Jan 3;87(2):24. doi: 10.1007/s11538-024-01400-2.
The human immune system can recognize, attack, and eliminate cancer cells, but cancers can escape this immune surveillance. Variants of ecological predator-prey models can capture the dynamics of such cancer control mechanisms by adaptive immune system cells. These dynamical systems describe, e.g., tumor cell-effector T cell conjugation, immune cell activation, cancer cell killing, and T cell exhaustion. Target (tumor) cell-T cell conjugation is integral to the adaptive immune system's cancer control and immunotherapy. However, whether conjugate dynamics should be explicitly included in mathematical models of cancer-immune interactions is incompletely understood. Here, we analyze the dynamics of a cancer-effector T cell system and focus on the impact of explicitly modeling the conjugate compartment to investigate the role of cellular conjugate dynamics. We formulate a deterministic modeling framework to compare possible equilibria and their stability, such as tumor extinction, tumor-immune coexistence (tumor control), or tumor escape. We also formulate the stochastic analog of this system to analyze the impact of demographic fluctuations that arise when cell populations are small. We find that explicit consideration of a conjugate compartment can (i) change long-term steady-state, (ii) critically change the time to reach an equilibrium, (iii) alter the probability of tumor escape, and (iv) lead to very different extinction time distributions. Thus, we demonstrate the importance of the conjugate compartment in defining tumor-effector T cell interactions. Accounting for transitionary compartments of cellular interactions may better capture the dynamics of tumor control and progression.

47Spatial interactions modulate tumor growth and immune infiltration.PubMed

Sadegh Marzban, Sonal Srivastava, Sharon Kartika, et al.
NPJ Syst Biol Appl. 2024 Sep 30;10(1):106. doi: 10.1038/s41540-024-00438-1.
Direct observation of tumor-immune interactions is unlikely in tumors with currently available technology, but computational simulations based on clinical data can provide insight to test hypotheses. It is hypothesized that patterns of collagen evolve as a mechanism of immune escape, but the exact nature of immune-collagen interactions is poorly understood. Spatial data quantifying collagen fiber alignment in squamous cell carcinomas indicates that late-stage disease is associated with highly aligned fibers. Our computational modeling framework discriminates between two hypotheses: immune cell migration that moves (1) parallel or (2) perpendicular to collagen fiber orientation. The modeling recapitulates immune-extracellular matrix interactions where collagen patterns provide immune protection, leading to an emergent inverse relationship between disease stage and immune coverage. Here, computational modeling provides important mechanistic insights by defining a kernel cell-cell interaction function that considers a spectrum of local (cell-scale) to global (tumor-scale) spatial interactions. Short-range interaction kernels provide a mechanism for tumor cell survival under conditions with strong Allee effects, while asymmetric tumor-immune interaction kernels lead to poor immune response. Thus, the length scale of tumor-immune interaction kernels drives tumor growth and infiltration.

48Solving neurodegeneration: common mechanisms and strategies for new treatments.PubMed

Lauren K Wareham, Shane A Liddelow, Sally Temple, et al.
Mol Neurodegener. 2022 Mar 21;17(1):23. doi: 10.1186/s13024-022-00524-0.
Across neurodegenerative diseases, common mechanisms may reveal novel therapeutic targets based on neuronal protection, repair, or regeneration, independent of etiology or site of disease pathology. To address these mechanisms and discuss emerging treatments, in April, 2021, Glaucoma Research Foundation, BrightFocus Foundation, and the Melza M. and Frank Theodore Barr Foundation collaborated to bring together key opinion leaders and experts in the field of neurodegenerative disease for a virtual meeting titled "Solving Neurodegeneration". This "think-tank" style meeting focused on uncovering common mechanistic roots of neurodegenerative disease and promising targets for new treatments, catalyzed by the goal of finding new treatments for glaucoma, the world's leading cause of irreversible blindness and the common interest of the three hosting foundations. Glaucoma, which causes vision loss through degeneration of the optic nerve, likely shares early cellular and molecular events with other neurodegenerative diseases of the central nervous system. Here we discuss major areas of mechanistic overlap between neurodegenerative diseases of the central nervous system: neuroinflammation, bioenergetics and metabolism, genetic contributions, and neurovascular interactions. We summarize important discussion points with emphasis on the research areas that are most innovative and promising in the treatment of neurodegeneration yet require further development. The research that is highlighted provides unique opportunities for collaboration that will lead to efforts in preventing neurodegeneration and ultimately vision loss.

49AI-driven discovery of brain-penetrant Galectin-3 inhibitors for Alzheimer's disease therapy.PubMed

Xueyan Liu, Jiexin Xu, Shuping Zheng, et al.
Pharmacol Res. 2025 Aug;218:107834. doi: 10.1016/j.phrs.2025.107834. Epub 2025 Jun 19.
Galectin-3 (Gal-3) has emerged as a critical regulator of neuroinflammation and a promising therapeutic target for Alzheimer's disease (AD). Nevertheless, the development of brain-penetrant small-molecule Gal-3 inhibitors poses a significant challenge. To address this, we employed an artificial intelligence (AI)-driven drug discovery platform, identifying FJMU1887 as a novel Gal-3 inhibitor possessing optimized pharmacokinetic properties and favorable blood-brain barrier (BBB) permeability. Following AI-based virtual screening and structure prioritization, FJMU1887 demonstrated direct binding to Gal-3 with an affinity (Kd) of 1.55 μM, determined by microscale thermophoresis (MST). Crucially, mechanistic studies revealed that FJMU1887 disrupts the Gal-3-TREM2 interaction, as evidenced by fluorescence resonance energy transfer (FRET) and fluorescence correlation spectroscopy (FCS) assays. In vitro, FJMU1887 suppressed inflammatory responses in BV-2 microglial cells, inhibiting TNF-α with an IC₅₀ of 2.36 ± 0.37 μM, without inducing cytotoxicity. Pharmacokinetic assessments via parallel artificial membrane permeability assay for BBB (PAMPA-BBB) and in situ brain perfusion revealed effective blood-brain barrier penetration by FJMU1887, though partial P-glycoprotein-mediated efflux was observed. In vivo, 30-day oral administration of FJMU1887 to 14-month-old 5 ×FAD mice significantly reduced Gal-3 expression, attenuated microglial activation and neuroinflammation, decreased amyloid-β burden, and restored synaptic integrity. Notably, FJMU1887 improved cognitive performance in both 5 ×FAD and oligomeric Aβ-induced cognitive impairment mouse models across multiple behavioral paradigms. Collectively, FJMU1887 represents a brain-penetrant small-molecule Gal-3 inhibitor with dual anti-neuroinflammatory and cognition-enhancing effects, establishing it as a promising lead compound for AD therapy.

50Z-DNA-binding protein 1 exacerbates myocardial ischemia‒reperfusion injury by inducing noncanonical cardiomyocyte PANoptosis.PubMed

Xiaokai Zhang, Shuai Song, Zihang Huang, et al.
Signal Transduct Target Ther. 2025 Oct 7;10(1):333. doi: 10.1038/s41392-025-02430-5.
Myocardial ischemia‒reperfusion (I/R) injury is the primary factor that counteracts the beneficial effects of reperfusion therapy. Cardiomyocyte death serves as the fundamental pathological hallmark of I/R injury. However, targeting a single type of cell death has been reported to be ineffective at preventing I/R injury. ZBP1 is well established as a nucleic acid sensor that activates inflammatory and various cell death signaling pathways. However, the specific role of ZBP1 in adult cardiomyocytes, particularly in the absence of nucleic acid ligands, remains largely unexplored. In this study, our dynamic transcriptomic analyses at various I/R stages revealed a cluster of genes significantly enriched in cell death-related processes, with ZBP1 showing significant expression changes in both our I/R injury mouse model and public human ischemic cardiomyopathy datasets. Cardiomyocytes are the primary cell type expressing ZBP1 in response to I/R injury. Hypoxia/reoxygenation stress induced the upregulation of multiple cell death markers indicative of PANoptosis in adult cardiomyocytes, which was mitigated by ZBP1 deficiency. Compared with treatment with conventional cell death inhibitors, cardiomyocyte-specific Zbp1 deficiency ameliorated I/R-induced PANoptosis, resulting in a more substantial reduction in myocardial infarct size. Conversely, myocardial Zbp1 overexpression in adult mice directly induced cardiac remodeling and heart failure. Mechanistically, ZBP1 drives cardiomyocyte PANoptosis by promoting the formation of the ZBP1/RIPK3/CAS8/CAS6 PANoptosome complex. Virtual screening and experimental validation revealed a novel small-molecule compound, MSB, which has high binding affinity for ZBP1 and effectively attenuates myocardial I/R injury both in vitro and in vivo. Collectively, these findings highlight the role of ZBP1 as a mediator of cardiomyocyte PANoptosis and suggest that targeting ZBP1 could be a promising strategy for mitigating myocardial I/R injury.

51Identification of a novel FOXO3 agonist that protects against alcohol induced liver injury.PubMed

Jinying Peng, Gaoshuang Liang, Yaqi Li, et al.
Biochem Biophys Res Commun. 2024 Apr 16;704:149690. doi: 10.1016/j.bbrc.2024.149690. Epub 2024 Feb 17.
Alcohol-related liver disease (ALD) is a global healthcare concern which caused by excessive alcohol consumption with limited treatment options. The pathogenesis of ALD is complex and involves in hepatocyte damage, hepatic inflammation, increased gut permeability and microbiome dysbiosis. FOXO3 is a well-recognized transcription factor which associated with longevity via promoting antioxidant stress response, preventing senescence and cell death, and inhibiting inflammation. We and many others have reported that FOXO3 mice develop more severe liver injury in response to alcohol. In the present study, we aimed to develop compounds that activate FOXO3 and further investigate their effects in alcohol induced liver injury. Through virtual screening, we discovered series of small molecular compounds that showed high affinity to FOXO3. We confirmed effects of compounds on FOXO3 target gene expression, as well as antioxidant and anti-apoptotic effects in vitro. Subsequently we evaluated the protective efficacy of compounds in alcohol induced liver injury in vivo. As a result, the leading compound we identified, 214991, activated downstream target genes expression of FOXO3, inhibited intracellular ROS accumulation and cell apoptosis induced by HO and sorafenib. By using Lieber-DeCarli alcohol feeding mouse model, 214991 showed protective effects against alcohol-induced liver inflammation, macrophage and neutrophil infiltration, and steatosis. These findings not only reinforce the potential of FOXO3 as a valuable target for therapeutic intervention of ALD, but also suggested that compound 214991 as a promising candidate for the development of innovative therapeutic strategies of ALD.

52The critical role of neutrophil-endothelial cell interactions in sepsis: new synergistic approaches employing organ-on-chip, omics, immune cell phenotyping and modeling to identify new therapeutics.PubMed

Dan Liu, Jordan C Langston, Balabhaskar Prabhakarpandian, et al.
Front Cell Infect Microbiol. 2024 Jan 8;13:1274842. doi: 10.3389/fcimb.2023.1274842. eCollection 2023.
Sepsis is a global health concern accounting for more than 1 in 5 deaths worldwide. Sepsis is now defined as life-threatening organ dysfunction caused by a dysregulated host response to infection. Sepsis can develop from bacterial (gram negative or gram positive), fungal or viral (such as COVID) infections. However, therapeutics developed in animal models and traditional sepsis models have had little success in clinical trials, as these models have failed to fully replicate the underlying pathophysiology and heterogeneity of the disease. The current understanding is that the host response to sepsis is highly diverse among patients, and this heterogeneity impacts immune function and response to infection. Phenotyping immune function and classifying sepsis patients into specific endotypes is needed to develop a personalized treatment approach. Neutrophil-endothelium interactions play a critical role in sepsis progression, and increased neutrophil influx and endothelial barrier disruption have important roles in the early course of organ damage. Understanding the mechanism of neutrophil-endothelium interactions and how immune function impacts this interaction can help us better manage the disease and lead to the discovery of new diagnostic and prognosis tools for effective treatments. In this review, we will discuss the latest research exploring how modeling of a synergistic combination of new organ-on-chip models incorporating human cells/tissue, omics analysis and clinical data from sepsis patients will allow us to identify relevant signaling pathways and characterize specific immune phenotypes in patients. Emerging technologies such as machine learning can then be leveraged to identify druggable therapeutic targets and relate them to immune phenotypes and underlying infectious agents. This synergistic approach can lead to the development of new therapeutics and the identification of FDA approved drugs that can be repurposed for the treatment of sepsis.

53Molecular mechanisms of epithelial-mesenchymal transition.PubMed

Samy Lamouille, Jian Xu, Rik Derynck
Nat Rev Mol Cell Biol. 2014 Mar;15(3):178-96. doi: 10.1038/nrm3758.
The transdifferentiation of epithelial cells into motile mesenchymal cells, a process known as epithelial-mesenchymal transition (EMT), is integral in development, wound healing and stem cell behaviour, and contributes pathologically to fibrosis and cancer progression. This switch in cell differentiation and behaviour is mediated by key transcription factors, including SNAIL, zinc-finger E-box-binding (ZEB) and basic helix-loop-helix transcription factors, the functions of which are finely regulated at the transcriptional, translational and post-translational levels. The reprogramming of gene expression during EMT, as well as non-transcriptional changes, are initiated and controlled by signalling pathways that respond to extracellular cues. Among these, transforming growth factor-β (TGFβ) family signalling has a predominant role; however, the convergence of signalling pathways is essential for EMT.

54EMT: Mechanisms and therapeutic implications.PubMed

Mohini Singh, Nicolas Yelle, Chitra Venugopal, et al.
Pharmacol Ther. 2018 Feb;182:80-94. doi: 10.1016/j.pharmthera.2017.08.009. Epub 2017 Aug 20.
Metastasis, the dissemination of cancer cells from primary tumors, represents a major hurdle in the treatment of cancer. The epithelial-mesenchymal transition (EMT) has been studied in normal mammalian development for decades, and it has been proposed as a critical mechanism during cancer progression and metastasis. EMT is tightly regulated by several internal and external cues that orchestrate the shifting from an epithelial-like phenotype into a mesenchymal phenotype, relying on a delicate balance between these two stages to promote metastatic development. EMT is thought to be induced in a subset of metastatic cancer stem cells (MCSCs), bestowing this population with the ability to spread throughout the body and contributing to therapy resistance. The EMT pathway is of increasing interest as a novel therapeutic avenue in the treatment of cancer, and could be targeted to prevent tumor cell dissemination in early stage patients or to eradicate existing metastatic cells in advanced stages. In this review, we describe the sequence of events and defining mechanisms that take place during EMT, and how these interactions drive cancer cell progression into metastasis. We summarize clinical interventions focused on targeting various aspects of EMT and their contribution to preventing cancer dissemination.

55EMT, MET, Plasticity, and Tumor Metastasis.PubMed

Basil Bakir, Anna M Chiarella, Jason R Pitarresi, et al.
Trends Cell Biol. 2020 Oct;30(10):764-776. doi: 10.1016/j.tcb.2020.07.003. Epub 2020 Aug 13.
Cancer cell identity and plasticity are required in transition states, such as epithelial-mesenchymal transition (EMT) and mesenchymal-epithelial transition (MET), in primary tumor initiation, progression, and metastasis. The functional roles of EMT, MET, and the partial state (referred to as pEMT) may vary based on the type of tumor, the state of dissemination, and the degree of metastatic colonization. Herein, we review EMT, MET, pEMT, and plasticity in the context of tumor metastasis.

56Epithelial-mesenchymal transition in tumor metastasis.PubMed

Kay T Yeung, Jing Yang
Mol Oncol. 2017 Jan;11(1):28-39. doi: 10.1002/1878-0261.12017. Epub 2016 Dec 9.
The epithelial-mesenchymal transition (EMT) is a developmental program that enables stationary epithelial cells to gain the ability to migrate and invade as single cells. Tumor cells reactivate EMT to acquire molecular alterations that enable the partial loss of epithelial features and partial gain of a mesenchymal phenotype. Our understanding of the contribution of EMT to tumor invasion, migration, and metastatic outgrowth has evolved over the past decade. In this review, we provide a summary of both historic and recent studies on the role of EMT in the metastatic cascade from various experimental systems, including cancer cell lines, genetic mouse tumor models, and clinical human breast cancer tissues.

57Non-redundant functions of EMT transcription factors.PubMed

Marc P Stemmler, Rebecca L Eccles, Simone Brabletz, et al.
Nat Cell Biol. 2019 Jan;21(1):102-112. doi: 10.1038/s41556-018-0196-y. Epub 2019 Jan 2.
Epithelial-mesenchymal transition (EMT) is a crucial embryonic programme that is executed by various EMT transcription factors (EMT-TFs) and is aberrantly activated in cancer and other diseases. However, the causal role of EMT and EMT-TFs in different disease processes, especially cancer and metastasis, continues to be debated. In this Review, we identify and describe specific, non-redundant functions of the different EMT-TFs and discuss the reasons that may underlie disputes about EMT in cancer.

58How to Simulate a Germinal Center.PubMed

Philippe A Robert, Ananya Rastogi, Sebastian C Binder, et al.
Methods Mol Biol. 2017;1623:303-334. doi: 10.1007/978-1-4939-7095-7_22.
Germinal centers host a mini-evolutionary environment where B cells can mutate their receptor and be selected depending on its affinity to target antigens in a process called affinity maturation. Starting from founder cells with a weak B cell receptor affinity, germinal centers release output cells as antibody-secreting cells or memory cells with a very high affinity, a property which is essential for pathogen clearance and immune memory. Therapeutic interventions on the germinal centers are tantalizing approaches to improve vaccines or to support rejection of chronic pathogens such as HIV. However, the complexity of the selection processes makes it very hard to make reliable predictions. Here, we present in detail how to build an agent-based model (hyphasma), accounting for the dynamics of the germinal center. It encompasses the core quantitative traits of affinity maturation, and allowed to make reliable predictions in previous studies.

59Current state and perspectives in modeling and control of human pluripotent stem cell expansion processes in stirred-tank bioreactors.PubMed

Vytautas Galvanauskas, Vykantas Grincas, Rimvydas Simutis, et al.
Biotechnol Prog. 2017 Mar;33(2):355-364. doi: 10.1002/btpr.2431. Epub 2017 Jan 10.
Implementation of model-based practices for process development, control, automation, standardization, and validation are important factors for therapeutic and industrial applications of human pluripotent stem cells. As robust cultivation strategies for pluripotent stem cell expansion and differentiation have yet to be determined, process development could be enhanced by application of mathematical models and advanced control systems to optimize growth conditions. Therefore, it is important to understand both the potential of possible applications and the apparent limitations of existing mathematical models to improve pluripotent stem cell cultivation technologies. In the present review, the authors focus on these issues as they apply to stem cell expansion processes. © 2017 American Institute of Chemical Engineers Biotechnol. Prog., 33:355-364, 2017.

60Interlinked bi-stable switches govern the cell fate commitment of embryonic stem cells.PubMed

Amitava Giri, Sandip Kar
FEBS Lett. 2024 Apr;598(8):915-934. doi: 10.1002/1873-3468.14832. Epub 2024 Feb 26.
The development of embryonic stem (ES) cells to extraembryonic trophectoderm and primitive endoderm lineages manifests distinct steady-state expression patterns of two key transcription factors-Oct4 and Nanog. How dynamically such kind of steady-state expressions are maintained remains elusive. Herein, we demonstrate that steady-state dynamics involving two bistable switches which are interlinked via a stepwise (Oct4) and a mushroom-like (Nanog) manner orchestrate the fate specification of ES cells. Our hypothesis qualitatively reconciles various experimental observations and elucidates how different feedback and feedforward motifs orchestrate the extraembryonic development and stemness maintenance of ES cells. Importantly, the model predicts strategies to optimize the dynamics of self-renewal and differentiation of embryonic stem cells that may have therapeutic relevance in the future.

61Mathematical modeling of regenerative processes.PubMed

Osvaldo Chara, Elly M Tanaka, Lutz Brusch
Curr Top Dev Biol. 2014;108:283-317. doi: 10.1016/B978-0-12-391498-9.00011-5.
In many animals, regenerative processes can replace lost body parts. Organ and tissue regeneration consequently also hold great medical promise. The regulation of regenerative processes is achieved through concerted actions of multiple organizational levels of the organism, from diffusing molecules and cellular gene expression patterns up to tissue mechanics. Our intuition is usually not adapted well to this degree of complexity and the quantitative aspects of the regulation of regenerative processes remain poorly understood. One way out of this dilemma lies in the combination of experimentation and mathematical modeling within an iterative process of model development/refinement, model predictions for novel experimental conditions, quantitative experiments testing these predictions, and subsequent model refinement. This interdisciplinary approach has already provided key insights into smaller scale processes during embryonic development and a so-far limited number of more complex regeneration processes. This review discusses selected theoretical and interdisciplinary studies and is structured along the three phases of regeneration: (1) initiation of a regeneration response, (2) tissue patterning during regenerate growth, (3) arresting the regeneration response. Moreover, we highlight the opportunities provided by extensions of mathematical models from developmental processes toward the study of related regenerative processes.

62Human Organoids: Tools for Understanding Biology and Treating Diseases.PubMed

Frans Schutgens, Hans Clevers
Annu Rev Pathol. 2020 Jan 24;15:211-234. doi: 10.1146/annurev-pathmechdis-012419-032611. Epub 2019 Sep 24.
Organoids are in vitro-cultured three-dimensional structures that recapitulate key aspects of in vivo organs. They can be established from pluripotent stem cells and from adult stem cells, the latter being the subject of this review. Organoids derived from adult stem cells exploit the tissue regeneration process that is driven by these cells, and they can be established directly from the healthy or diseased epithelium of many organs. Organoids are amenable to any experimental approach that has been developed for cell lines. Applications in experimental biology involve the modeling of tissue physiology and disease, including malignant, hereditary, and infectious diseases. Biobanks of patient-derived tumor organoids are used in drug development research, and they hold promise for developing personalized and regenerative medicine. In this review, we discuss the applications of adult stem cell-derived organoids in the laboratory and the clinic.

63Reversed graph embedding resolves complex single-cell trajectories.PubMed

Xiaojie Qiu, Qi Mao, Ying Tang, et al.
Nat Methods. 2017 Oct;14(10):979-982. doi: 10.1038/nmeth.4402. Epub 2017 Aug 21.
Single-cell trajectories can unveil how gene regulation governs cell fate decisions. However, learning the structure of complex trajectories with multiple branches remains a challenging computational problem. We present Monocle 2, an algorithm that uses reversed graph embedding to describe multiple fate decisions in a fully unsupervised manner. We applied Monocle 2 to two studies of blood development and found that mutations in the genes encoding key lineage transcription factors divert cells to alternative fates.

64Phylodynamics for cell biologists.PubMed

T Stadler, O G Pybus, M P H Stumpf
Science. 2021 Jan 15;371(6526). doi: 10.1126/science.aah6266.
Multicellular organisms are composed of cells connected by ancestry and descent from progenitor cells. The dynamics of cell birth, death, and inheritance within an organism give rise to the fundamental processes of development, differentiation, and cancer. Technical advances in molecular biology now allow us to study cellular composition, ancestry, and evolution at the resolution of individual cells within an organism or tissue. Here, we take a phylogenetic and phylodynamic approach to single-cell biology. We explain how "tree thinking" is important to the interpretation of the growing body of cell-level data and how ecological null models can benefit statistical hypothesis testing. Experimental progress in cell biology should be accompanied by theoretical developments if we are to exploit fully the dynamical information in single-cell data.

65Virtual Cell Challenge: Toward a Turing test for the virtual cell.PubMed

Yusuf H Roohani, Tony J Hua, Po-Yuan Tung, et al.
Cell. 2025 Jun 26;188(13):3370-3374. doi: 10.1016/j.cell.2025.06.008.
Virtual cells are an emerging frontier at the intersection of artificial intelligence and biology. A key goal of these cell state models is predicting cellular responses to perturbations. The Virtual Cell Challenge is being established to catalyze progress toward this goal. This recurring and open benchmark competition from the Arc Institute will provide an evaluation framework, purpose-built datasets, and a venue for accelerating model development.

66Mathematical Model-Driven Deep Learning Enables Personalized Adaptive Therapy.PubMed

Kit Gallagher, Maximilian A R Strobl, Derek S Park, et al.
Cancer Res. 2024 Jun 4;84(11):1929-1941. doi: 10.1158/0008-5472.CAN-23-2040.
UNLABELLED: Standard-of-care treatment regimens have long been designed for maximal cell killing, yet these strategies often fail when applied to metastatic cancers due to the emergence of drug resistance. Adaptive treatment strategies have been developed as an alternative approach, dynamically adjusting treatment to suppress the growth of treatment-resistant populations and thereby delay, or even prevent, tumor progression. Promising clinical results in prostate cancer indicate the potential to optimize adaptive treatment protocols. Here, we applied deep reinforcement learning (DRL) to guide adaptive drug scheduling and demonstrated that these treatment schedules can outperform the current adaptive protocols in a mathematical model calibrated to prostate cancer dynamics, more than doubling the time to progression. The DRL strategies were robust to patient variability, including both tumor dynamics and clinical monitoring schedules. The DRL framework could produce interpretable, adaptive strategies based on a single tumor burden threshold, replicating and informing optimal treatment strategies. The DRL framework had no knowledge of the underlying mathematical tumor model, demonstrating the capability of DRL to help develop treatment strategies in novel or complex settings. Finally, a proposed five-step pathway, which combined mechanistic modeling with the DRL framework and integrated conventional tools to improve interpretability compared with traditional "black-box" DRL models, could allow translation of this approach to the clinic. Overall, the proposed framework generated personalized treatment schedules that consistently outperformed clinical standard-of-care protocols. SIGNIFICANCE: Generation of interpretable and personalized adaptive treatment schedules using a deep reinforcement framework that interacts with a virtual patient model overcomes the limitations of standardized strategies caused by heterogeneous treatment responses.

67CNGBdb: China National GeneBank DataBase.PubMed

Feng Zhen Chen, Li Jin You, Fan Yang, et al.
Yi Chuan. 2020 Aug 20;42(8):799-809. doi: 10.16288/j.yczz.20-080.
China National GeneBank DataBase (CNGBdb) is a data platform aiming to systematically archiving and sharing of multi-omics data in life science. As the service portal of Bio-informatics Data Center of the core structure, namely, "Three Banks and Two Platforms" of China National GeneBank (CNGB), CNGBdb has the advantages of rich sample resources, data resources, cooperation projects, powerful data computation and analysis capabilities. With the advent of high throughput sequencing technologies, research in life science has entered the big data era, which is in the need of closer international cooperation and data sharing. With the development of China's economy and the increase of investment in life science research, we need to establish a national public platform for data archiving and sharing in life science to promote the systematic management, application and industrial utilization. Currently, CNGBdb can provide genomic data archiving, information search engines, data management and data analysis services. The data schema of CNGBdb has covered projects, samples, experiments, runs, assemblies, variations and sequences. Until May 22, 2020, CNGBdb has archived 2176 research projects and more than 2221 TB sequencing data submitted by researchers globally. In the future, CNGBdb will continue to be dedicated to promoting data sharing in life science research and improving the service capability. CNGBdb website is: https://db.cngb.org/.

68PhenoDB, GeneMatcher and VariantMatcher, tools for analysis and sharing of sequence data.PubMed

Elizabeth Wohler, Renan Martin, Sean Griffith, et al.
Orphanet J Rare Dis. 2021 Aug 18;16(1):365. doi: 10.1186/s13023-021-01916-z.
BACKGROUND: With the advent of whole exome (ES) and genome sequencing (GS) as tools for disease gene discovery, rare variant filtering, prioritization and data sharing have become essential components of the search for disease genes and variants potentially contributing to disease phenotypes. The computational storage, data manipulation, and bioinformatic interpretation of thousands to millions of variants identified in ES and GS, respectively, is a challenging task. To aid in that endeavor, we constructed PhenoDB, GeneMatcher and VariantMatcher. RESULTS: PhenoDB is an accessible, freely available, web-based platform that allows users to store, share, analyze and interpret their patients' phenotypes and variants from ES/GS data. GeneMatcher is accessible to all stakeholders as a web-based tool developed to connect individuals (researchers, clinicians, health care providers and patients) around the globe with interest in the same gene(s), variant(s) or phenotype(s). Finally, VariantMatcher was developed to enable public sharing of variant-level data and phenotypic information from individuals sequenced as part of multiple disease gene discovery projects. Here we provide updates on PhenoDB and GeneMatcher applications and implementation and introduce VariantMatcher. CONCLUSION: Each of these tools has facilitated worldwide data sharing and data analysis and improved our ability to connect genes to phenotypic traits. Further development of these platforms will expand variant analysis, interpretation, novel disease-gene discovery and facilitate functional annotation of the human genome for clinical genomics implementation and the precision medicine initiative.

69The ViReflow pipeline enables user friendly large scale viral consensus genome reconstruction.PubMed

Niema Moshiri, Kathleen M Fisch, Amanda Birmingham, et al.
Sci Rep. 2022 Mar 24;12(1):5077. doi: 10.1038/s41598-022-09035-w.
Throughout the COVID-19 pandemic, massive sequencing and data sharing efforts enabled the real-time surveillance of novel SARS-CoV-2 strains throughout the world, the results of which provided public health officials with actionable information to prevent the spread of the virus. However, with great sequencing comes great computation, and while cloud computing platforms bring high-performance computing directly into the hands of all who seek it, optimal design and configuration of a cloud compute cluster requires significant system administration expertise. We developed ViReflow, a user-friendly viral consensus sequence reconstruction pipeline enabling rapid analysis of viral sequence datasets leveraging Amazon Web Services (AWS) cloud compute resources and the Reflow system. ViReflow was developed specifically in response to the COVID-19 pandemic, but it is general to any viral pathogen. Importantly, when utilized with sufficient compute resources, ViReflow can trim, map, call variants, and call consensus sequences from amplicon sequence data from 1000 SARS-CoV-2 samples at 1000X depth in < 10 min, with no user intervention. ViReflow's simplicity, flexibility, and scalability make it an ideal tool for viral molecular epidemiological efforts.

70Discovery of senolytics using machine learning.PubMed

Vanessa Smer-Barreto, Andrea Quintanilla, Richard J R Elliott, et al.
Nat Commun. 2023 Jun 10;14(1):3445. doi: 10.1038/s41467-023-39120-1.
Cellular senescence is a stress response involved in ageing and diverse disease processes including cancer, type-2 diabetes, osteoarthritis and viral infection. Despite growing interest in targeted elimination of senescent cells, only few senolytics are known due to the lack of well-characterised molecular targets. Here, we report the discovery of three senolytics using cost-effective machine learning algorithms trained solely on published data. We computationally screened various chemical libraries and validated the senolytic action of ginkgetin, periplocin and oleandrin in human cell lines under various modalities of senescence. The compounds have potency comparable to known senolytics, and we show that oleandrin has improved potency over its target as compared to best-in-class alternatives. Our approach led to several hundred-fold reduction in drug screening costs and demonstrates that artificial intelligence can take maximum advantage of small and heterogeneous drug screening data, paving the way for new open science approaches to early-stage drug discovery.