• Suppr超能文献
  • 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
定价套餐&价格
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Win 客户端微信小程序
定价
会员套餐积分包API 积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026
  1. 首页
  2. 分享广场
  3. 晚期肺癌免疫治疗疗效预测模型研究进展与多参数建模方法解析

晚期肺癌免疫治疗疗效预测模型研究进展与多参数建模方法解析

深度研究匿名用户发表于 2025年10月31日 09:355阅读
发起深度研究
发起深度研究

1. 晚期肺癌免疫治疗的临床背景与预测需求

1.1 免疫治疗在晚期肺癌中的应用现状

晚期非小细胞肺癌(NSCLC)的治疗在过去十年中取得了显著进展,特别是免疫检查点抑制剂(ICIs)的出现,极大地改善了患者的生存预后 1。以程序性死亡受体-1(PD-1)及其配体(PD-L1)抑制剂为代表的免疫疗法,通过激活人体自身免疫系统来对抗肿瘤,已成为晚期NSCLC一线治疗的重要组成部分,甚至对于不携带可靶向致癌基因的晚期NSCLC患者,PD-1或PD-L1抑制剂治疗已几乎成为常规 12。这类疗法在部分患者中实现了前所未有的长期生存,为以往预后极差的患者带来了新的希望 1。

然而,尽管免疫治疗取得了突破性成功,其疗效在不同患者之间存在显著的异质性 2。并非所有患者都能从免疫治疗中获益,部分患者甚至可能出现原发性耐药或在治疗过程中获得性耐药 3。例如,针对晚期NSCLC的纳武利尤单抗(Nivolumab)或帕博利珠单抗(Pembrolizumab)等PD-1/PD-L1抑制剂,虽然在临床试验中显示出积极疗效,但整体响应率仍有提升空间 4。这种疗效上的不确定性凸显了对个体化预测模型的需求,旨在精准识别最有可能从免疫治疗中获益的患者,从而优化治疗方案,避免不必要的毒副作用和医疗资源浪费。现有研究表明,即便在PD-L1表达低于1%的NSCLC患者中,新辅助化疗免疫疗法相较于单纯化疗也能带来事件-无进展生存期(EFS)的显著改善,进一步强调了预测模型在精细化患者分层中的潜力 5。

1.2 疗效预测的核心临床价值

在晚期肺癌免疫治疗中,构建精准的疗效预测模型具有核心临床价值。首先,它能够实现治疗方案的优化。目前,免疫治疗药物可能伴随免疫相关不良事件(irAEs),虽然大部分irAEs可控,但仍有潜在风险。通过预测模型,医生可以在治疗前筛选出最有可能从免疫治疗中获益的患者,避免将无效治疗强加给不适宜的患者,从而将有限的医疗资源分配给最需要的群体 6。例如,有研究指出,通过深度学习模型预测非小细胞肺癌患者免疫治疗后的无进展生存期(PFS),能够有效区分预后良好和不良的患者,从而指导个性化治疗 7。

其次,疗效预测模型有助于显著降低无效治疗的风险。对于那些对免疫治疗响应不佳的患者,如果能提前识别,可以避免其承受不必要的治疗毒性、经济负担以及延误其他可能有效的治疗时机。例如,在晚期膀胱癌的研究中,预测模型能够整合突变数据和基因表达数据,识别与免疫检查点抑制剂响应相关的关键因素,从而为精准医疗提供依据 8。对于小细胞肺癌(SCLC),结合临床特征和影像组学特征的整合模型,也能有效预测化疗免疫疗法的治疗效果,区分高进展风险患者,避免其接受无效治疗 9。

最后,精准的疗效预测模型对提升患者生存质量和生活预后具有深远意义。通过选择最合适的治疗策略,患者可以避免无效治疗带来的身体和心理负担,减少并发症,并有可能延长高质量的生存期。例如,通过对患者预处理CT影像进行深度学习分析,可以独立于现有临床病理生物标志物提供预测信息,指导精准免疫治疗,进一步改善患者预后 10。此外,通过预测模型识别治疗中的不良事件,也可以提前采取预防措施,例如在急性淋巴细胞白血病(ALL)中,识别影响治疗决策的变量,甚至可以预测治疗后严重不良事件的发生,从而考虑适当的预防措施 11。这些都指向了以患者为中心的精准医疗方向,即根据每个患者的独特生物学特征,量身定制治疗方案,最大化治疗效果,最小化不良反应,最终全面提升患者的生存质量和整体预后。

2. 肺癌免疫治疗疗效预测模型的研究进展与关键影响因子

2.1 单参数预测模型的局限性与突破

在肺癌免疫治疗领域,早期的疗效预测主要依赖于单一生物标志物,其中最具代表性的是程序性死亡配体1(PD-L1)表达、肿瘤突变负荷(TMB)和微卫星不稳定性(MSI)。这些单一标志物在一定程度上展现了预测潜力,但其局限性也日益凸显。

PD-L1表达是通过免疫组织化学(IHC)检测肿瘤细胞或免疫细胞表面PD-L1蛋白的水平,被认为是预测PD-1/PD-L1抑制剂疗效的首要伴随诊断标志物 1213。高PD-L1表达通常预示着患者对免疫治疗有更好的响应和更长的无进展生存期(PFS)或总生存期(OS) 14。然而,PD-L1表达并非一个完美的预测因子。首先,PD-L1表达具有异质性和动态性,其水平可能因肿瘤内部不同区域、不同时间点以及治疗前后的变化而有所差异 1315。其次,部分PD-L1高表达的患者可能对免疫治疗无响应,而一些PD-L1低表达甚至阴性的患者也能从治疗中获益 1316。这表明PD-L1表达的预测效能存在“天花板”,单一检测难以全面反映复杂的肿瘤免疫微环境 13。

肿瘤突变负荷(TMB)是指肿瘤基因组中体细胞非同义突变的总数量,高TMB的肿瘤细胞被认为更容易产生新抗原,从而激活更强的抗肿瘤免疫反应 17。研究表明,高TMB与多种癌症(包括肺癌)患者对免疫检查点抑制剂的响应呈正相关 1718。特别是对于非小细胞肺癌,高TMB被认为是独立于PD-L1表达的有效预测标志物,能够预测更好的客观缓解率和更长的生存期 1218。然而,TMB的检测标准化仍面临挑战,包括不同的检测方法(如全外显子测序WES与靶向测序panels)和阈值设定可能导致结果差异 1317。此外,高TMB也并非万能,某些具有特定基因突变(如EGFR或ALK突变)的NSCLC患者,即使TMB较高,对免疫治疗的益处也有限 14。

微卫星不稳定性(MSI)是由于DNA错配修复(MMR)系统功能缺陷导致基因组中微卫星区域长度发生变化的一种遗传表型。MSI高(MSI-H)的肿瘤通常伴随高TMB,并且对免疫治疗表现出更强的敏感性 1219。MSI-H已成为多个瘤种(包括肺癌)的泛癌种免疫治疗响应标志物 1220。然而,MSI在肺癌中的发生率相对较低,这意味着它只能筛选出小部分可能从免疫治疗中获益的患者 1219。

总体而言,PD-L1表达、TMB和MSI作为单一预测标志物,尽管在临床实践中发挥了重要作用,但其各自的局限性使得精准预测仍面临挑战 1321。这些单一标志物无法完全捕捉肿瘤免疫微环境的复杂性和多变性,因此,探索多参数整合模型成为突破单一标志物预测效能瓶颈的关键方向。

2.2 多参数模型的核心影响因子整合

鉴于单一生物标志物在预测肺癌免疫治疗疗效上的局限性,当前研究正积极转向整合多维度参数的多参数预测模型。这种方法旨在全面捕捉肿瘤生物学特征、宿主免疫状态以及治疗过程中的动态变化,从而更精准地预测患者的治疗响应和生存获益。多参数模型的核心在于协同整合来自临床、组学和影像等不同来源的数据,以揭示更深层次的预测机制。

临床指标的整合:患者的临床特征是预测模型的重要组成部分。例如,体能状态(ECOG PS评分)直接反映了患者的全身健康状况和对治疗的耐受能力,ECOG PS评分越低(即体能状态越好)的患者通常预后更佳,对免疫治疗的响应也可能更好。肿瘤的病理分期、转移灶的数量和位置也对疗效预测有显著影响。研究发现,基线时的肝转移或多部位转移可能预示着对免疫治疗的响应率较低和预后不良。此外,血液学指标,如改良肺免疫预测指数(mLIPI)已被证明与免疫治疗的疗效相关 22。一项研究构建了结合年龄、基线间质性肺病、肺气肿与影像组学特征的联合模型,在预测免疫检查点抑制剂相关肺炎(CIP)的发生风险方面表现出良好准确性,提示临床特征与影像学数据结合的潜力 23。

组学数据的协同:随着高通量测序技术的发展,包括基因组学(如体细胞突变、拷贝数变异)、转录组学(基因表达谱)、蛋白组学以及微生物组学在内的组学数据为免疫治疗疗效预测提供了丰富的生物学信息。

  • 基因组学:除了TMB,特定基因突变(如EGFR、ALK等驱动基因突变)对免疫治疗疗效有显著影响。例如,EGFR突变型NSCLC患者对PD-1抑制剂的响应率通常较低,即便通过新辅助免疫化疗,EGFR突变型肺癌患者也可能表现出不同的免疫抵抗表型,提示需要更精细的分子分型来预测疗效 24。整合基因表达谱可以识别与免疫激活或抑制相关的通路,例如有研究通过分析肺腺癌(LUAD)中蛋白酪氨酸磷酸酶受体O型(PTPRO)的表达,发现其与患者预后及肿瘤免疫微环境(TIM)密切相关,可能成为潜在的免疫治疗靶点和预测因子 25。
  • 微生物组学:肠道微生物群被认为是影响免疫治疗疗效的重要因素。有研究基于非小细胞肺癌患者的宏基因组测序数据,构建了肠道菌群拓扑评分(TOPOSCORE),该评分结合了与免疫治疗抵抗和响应相关的特定菌种丰度,并在多个独立队列中得到验证,表明肠道菌群的生态拓扑结构可以作为预测癌症免疫治疗结果的动态诊断工具 26。

影像学特征的集成:医学影像(如CT、PET/CT)是临床常规检查,从中提取的影像组学特征(Radiomics)能够量化肿瘤内部的异质性,提供肿瘤微环境的非侵入性信息。

  • CT影像组学:通过分析治疗前CT图像中的纹理、形状和强度特征,可以构建预测模型。例如,一项研究从非小细胞肺癌患者治疗前CT图像中提取放射组学特征,并与临床病理特征相结合,成功预测了免疫检查点抑制剂的临床获益和无进展生存期 27。另一项研究则利用增强CT图像的delta-radiomics特征(治疗前后影像特征的变化)结合血液学指标,预测新辅助免疫化疗后的病理完全缓解(pCR),显示出较好的预测性能 28。
  • PET/CT影像组学:18F-FDG PET/CT能够反映肿瘤的代谢活性,其影像组学特征可以提供肿瘤的代谢异质性信息。有研究将18F-FDG PET/CT影像组学特征与临床数据相结合,成功构建了预测III期NSCLC患者生存率的列线图模型 29。此外,PET/CT影像组学模型也被用于预测新辅助免疫化疗后的病理完全缓解,其预测能力优于单一的PET或CT影像组学模型,且与肿瘤增殖抑制及抗肿瘤免疫细胞浸润相关联,揭示了潜在的生物学机制 30。

多参数协同预测的实例:一项针对NSCLC患者的研究整合了临床参数、影像组学特征和免疫特征数据,构建了一个多维预测列线图。该模型在预测PFS和OS方面均优于任何单一模型,其C指数和AUC值显著更高(PFS的AUC为0.771,OS的AUC为0.768) 22。这充分展示了多参数整合的优势,即不同维度的数据能够相互补充,共同提升模型的预测效能,从而为患者提供更加个性化和精准的治疗策略。这种整合方法不仅提高了预测的准确性,也为深入理解免疫治疗响应机制提供了新的视角。

3. “多参数建模”的技术内涵与方法学辨析

“多参数建模”在当前医学研究中通常指的是对来自不同维度和类型的数据进行整合,以构建更全面、更精准的预测或诊断模型。这种建模方法并非单一指代多模态数据、机器学习或大语言模型(LLMs),而是涵盖了这些技术在不同数据类型融合与分析中的应用。具体而言,它强调的是从多个信息源(如临床病理、影像、基因组学、蛋白质组学等)获取数据,并通过高级计算方法(如机器学习、深度学习)进行集成分析,以克服单一数据类型或单一标志物的局限性。

3.1 多模态数据融合的建模逻辑

多模态数据融合是“多参数建模”的核心策略之一,其基本逻辑在于将来源于不同模态(即不同类型和来源)的异质数据进行有效整合,以揭示数据内部更深层次的关联性和互补性,从而提升模型的预测能力。在晚期肺癌免疫治疗疗效预测中,常见的多模态数据包括临床数据(如年龄、性别、ECOG评分、病理类型、治疗史等)、组学数据(如基因测序、转录组、蛋白组等)和影像数据(如CT、PET/CT、MRI等)。这些数据模态各自提供了关于肿瘤和宿主状态的不同视角,通过融合可以形成更全面的患者表征。

异质特征整合方法:
多模态数据的整合并非简单的数据拼接,而是需要根据不同模态数据的特性,采用合适的融合策略。常见的方法包括:

  1. 特征级融合(Feature-level Fusion):在模型输入之前,从不同模态数据中提取有意义的特征,然后将这些特征向量进行拼接或加权组合,形成一个统一的特征向量,再输入到机器学习模型中进行训练。例如,从CT影像中提取放射组学特征,从基因测序数据中提取突变特征,然后将这些特征合并。
  2. 决策级融合(Decision-level Fusion):为每个模态数据训练一个独立的预测模型,然后将各个模型的预测结果(如概率、分类标签)进行组合(如投票、加权平均或元学习),得到最终的预测。这种方法可以保留各模态数据的独立性,但也可能忽略模态间的内在关联。
  3. 深度学习融合(Deep Learning Fusion):利用深度学习模型(如多头神经网络、卷积神经网络等)自动从不同模态数据中学习和提取特征,并在网络的不同层级进行融合。这种方法能够捕获更复杂的非线性关系和隐藏模式,是目前多模态数据融合研究的热点 3132333435。

PET-CT代谢参数联合循环肿瘤DNA(ctDNA)的实例:
PET/CT影像,特别是18F-FDG PET/CT,通过测量葡萄糖代谢水平来反映肿瘤的代谢活性,提供了肿瘤异质性和侵袭性的非侵入性信息。其代谢参数,如标准化摄取值(SUVmax, SUVmean)、代谢肿瘤体积(MTV)和总病灶糖酵解(TLG),已显示出与肿瘤的生物学特性和预后相关 36。例如,高MTV和TLG可能指示更具侵袭性的肿瘤表型。

循环肿瘤DNA(ctDNA)则是一种液体活检技术,通过分析外周血中肿瘤来源的DNA片段,能够实时、动态地反映肿瘤的基因组学信息,包括突变状态、疾病负荷和耐药机制。ctDNA具有微创、可重复取样、能够捕捉肿瘤异质性等优势,在肿瘤早期诊断、疗效监测、复发预测等方面展现出巨大潜力 37383940。

将PET-CT代谢参数与ctDNA数据进行融合,能够充分发挥两者的互补性,提供更全面的肿瘤信息:

  • PET-CT提供空间和代谢信息:PET-CT能够直观展示肿瘤的位置、大小、代谢活性及异质性,有助于评估肿瘤负荷和对治疗的早期反应。例如,一项研究利用PET/CT的影像特征结合ctDNA来预测非小细胞肺癌患者术后复发风险,发现PET/CT的“栖息地成像”(habitat imaging)亚型与ctDNA结合,在预测疾病复发方面具有互补价值 41。
  • ctDNA提供分子和动态信息:ctDNA能够检测微小残留病灶(MRD),揭示肿瘤的分子遗传学特征,并在治疗过程中提供实时的动态变化信息。例如,ctDNA水平的下降常预示着治疗有效,而其升高则可能提示疾病进展或复发 3840。
  • 跨模态信息互补性:PET-CT和ctDNA在不同层面对肿瘤进行刻画。PET-CT可捕获肿瘤的代谢异质性,而ctDNA则能反映肿瘤的遗传变异。两者结合可以实现优势互补,例如,PET/CT可能发现宏观病灶,但ctDNA能更早地捕获分子水平的微小残留病灶,或揭示导致PET/CT异常代谢背后的基因突变。对于结直肠癌,PET/CT的代谢参数已被证实能够预测微卫星不稳定性(MSI)状态,这表明影像学特征与分子生物学特征之间存在关联 36。这种互补性使得通过多模态融合能够构建出比单一模态更强大、更鲁棒的预测模型。例如,在乳腺癌中,结合MRI、病理全切片图像和临床风险因素的多模态系统,在预测新辅助化疗后的病理完全缓解(pCR)方面表现出优异的性能,显著优于单一模态模型 35。肺腺癌的生存预测模型也证实了多组学(基因表达、体细胞突变、临床数据)融合的优势 42。

这种融合策略不仅提升了预测的准确性,也为深入理解肿瘤生物学机制提供了新的视角,有助于实现更精准的个体化治疗。

3.2 机器学习与深度学习的模型选择

在多参数建模的框架下,机器学习(Machine Learning, ML)和深度学习(Deep Learning, DL)是实现数据融合和模式识别的核心技术。两者在处理复杂、高维数据方面各有侧重,并在肺癌免疫治疗疗效预测模型的构建中展现出不同的适用性和优势。

传统机器学习模型,如逻辑回归(Logistic Regression)、支持向量机(Support Vector Machine, SVM)、随机森林(Random Forest)和梯度提升树(Gradient Boosting Decision Tree, GBDT)等,在处理结构化数据(如临床指标、预处理后的组学特征)时表现出良好的性能。

  • 优势:这些模型通常具有较好的可解释性,模型训练所需数据量相对较小,且计算资源需求较低。对于特征工程良好、特征维度适中的数据集,传统机器学习模型能够快速建立预测模型,并识别出重要的预测因子。例如,在预测癌症预后时,基于临床和基因组数据的机器学习模型能够提供有效的风险分层 43。
  • 局限性:传统机器学习模型在处理非结构化数据(如原始影像、高维组学数据)时,需要繁琐的特征提取和选择过程。它们难以自动从原始数据中学习高级特征,且对于数据内部的复杂非线性关系建模能力有限。

深度学习模型,尤其是卷积神经网络(Convolutional Neural Network, CNN)和图神经网络(Graph Neural Network, GNN),则在处理大规模、高维、非结构化数据方面展现出强大能力。

  • 优势:深度学习模型能够自动从原始数据中学习和提取层次化的特征,无需人工干预,极大地简化了特征工程的复杂性。它们擅长捕捉数据内部的复杂非线性模式和潜在关联,尤其在图像识别和自然语言处理领域取得了突破性进展。在医疗领域,深度学习在医学图像分析(如肺结节检测和分类)、病理图像分析以及多组学数据融合等方面表现出色 4445464748。例如,MultiSurv模型利用多模态深度学习方法,整合临床、影像和高维组学数据,实现了对多种癌症的长期生存预测,并能处理缺失数据 49。另一个深度学习框架BioFusionNet也通过融合图像特征、遗传数据和临床数据,在乳腺癌生存风险分层中超越了现有先进方法 50。
  • 局限性:深度学习模型通常需要大量的标注数据进行训练,对计算资源(如GPU)要求较高。此外,其“黑箱”特性使得模型内部决策过程难以解释,这在临床应用中可能成为接受度的障碍。

肺结节影像-组学联合建模案例分析:
在肺癌领域,肺结节的早期诊断和良恶性判断是提高患者生存率的关键。传统的评估方法依赖于CT影像的形态学特征,但准确性仍有提升空间。结合影像组学和基因组学数据,深度学习模型展现出显著优势。

例如,一项研究利用深度学习技术,对肺结节CT影像进行分析,实现了较高的准确率 4651。进一步地,如果将肺结节的影像特征(通过CNN提取)与患者的基因突变信息(如EGFR、KRAS等)、基因表达谱数据以及临床病理数据进行融合,可以构建出更强大的预测模型。

在这样的联合建模中:

  • 影像数据处理:通常使用CNN来处理CT影像,自动学习和提取结节的深度特征,包括其形状、纹理、密度等非直观特征。这些特征比传统影像组学手工提取的特征更具判别力。
  • 组学数据处理:基因组学和转录组学数据可以预处理成结构化特征,或者通过特定的神经网络层(如全连接层、变分自编码器)进行嵌入,将其转换为与影像特征兼容的向量表示 50。
  • 数据融合策略:
    • 早期融合(Early Fusion):在特征提取阶段就将影像特征和组学特征进行拼接,然后输入到统一的深度学习网络中。
    • 晚期融合(Late Fusion):分别训练影像模型和组学模型,最后将各自模型的输出(如预测概率或高层特征)进行组合。
    • 中间融合(Intermediate Fusion)/混合融合(Hybrid Fusion):在深度网络的中间层进行不同模态特征的融合,允许模型在学习过程中逐步整合多模态信息。有研究提出,将临床数据在卷积操作之前与影像特征进行融合(pre-spatial fusion),能够显著提高模型的预测性能 5253。

通过深度学习实现影像-组学联合建模,能够更全面地捕捉肺结节的生物学行为和患者的个体差异,从而提高良恶性诊断的准确性,甚至预测其对免疫治疗的响应潜力。与传统机器学习相比,深度学习在处理原始医学影像和高维组学数据时的自动特征学习能力是其核心优势,使得模型能够发现更深层次的生物标志物和预测模式。尽管存在可解释性挑战,但通过可解释性AI(XAI)技术的发展,正在逐步解决这一问题,使其在临床转化中更具潜力 54。

3.3 大语言模型(LLMs)的潜在应用与边界

近年来,以GPT系列为代表的大语言模型(LLMs)在自然语言处理领域取得了突破性进展,其强大的文本生成、理解和推理能力为医学领域带来了新的应用前景。在多参数建模中,LLMs主要通过处理非结构化临床文本来挖掘潜在信息,辅助疗效预测 55。

LLMs在非结构化临床文本挖掘中的价值:
临床数据中存在大量非结构化文本,如病理报告、放射科报告、住院记录、门诊随访记录、手术记录以及多学科会诊(MDT)意见等。这些文本包含了患者详细的病史、症状描述、诊断依据、治疗过程、疗效评估及不良反应等关键信息。传统上,从这些文本中提取结构化信息需要大量人工阅读和编码,效率低下且易受主观性影响。LLMs的出现极大地提升了这一过程的自动化和智能化水平:

  1. 信息提取与结构化:LLMs能够自动识别并提取文本中的关键实体(如肿瘤类型、分期、基因突变、治疗方案、不良事件)及其相互关系,将其转化为结构化数据。例如,有研究表明LLMs能够有效地从放射科报告中提取关键临床数据,例如评估前列腺MRI报告中的放射学特征,展现出98.6%的平均准确率 5657。这对于构建预测模型所需的数据集至关重要。
  2. 语义理解与推理:LLMs不仅能提取字面信息,还能理解文本的深层语义,进行上下文推理。例如,它们可以分析病理报告中的描述性文字,判断肿瘤的侵袭性特征;或从随访记录中识别治疗效果的细微变化,甚至能够从临床总结中预测神经退行性疾病的病理诊断,其准确性已接近或达到初级放射科医生的水平 5859。在临床试验患者匹配方面,LLMs也被用于处理复杂的患者数据和入组排除标准,提高匹配效率 60。
  3. 弱标签生成:LLMs可以根据大量的非结构化报告,为图像或病例生成“弱标签”,这些标签可以作为训练其他机器学习模型的输入,减少人工标注的工作量 55。
  4. 辅助决策与报告生成:虽然不是直接用于预测,但LLMs能够生成总结报告、辅助医生进行鉴别诊断,甚至协助预测手术时长,从而提高临床工作效率 61。例如,在精神卫生领域,LLMs能够处理非结构化临床笔记,协助进行患者分诊,缩短等待时间 62。

LLMs在数值型/影像数据建模中的局限性:
尽管LLMs在文本处理方面表现出色,但其本质是基于自然语言的序列模型,这决定了它们在处理数值型数据和影像数据时存在固有的局限性:

  1. 数值型数据处理能力不足:LLMs不擅长直接进行复杂的数值计算、统计分析或处理结构化的表格数据(如实验室检查结果、生理参数等)。它们无法像专门的数值分析算法那样精确地识别数值之间的关系或进行趋势分析。虽然可以将数值数据转化为文本描述输入LLM,但这会导致信息损失,且依赖于文本描述的准确性。在预测败血症等任务中,LLMs需要从非结构化临床笔记中提取症状,但其核心的风险评分计算仍依赖于结构化数据和传统算法 63。
  2. 无法直接处理原始影像数据:LLMs无法直接“看懂”医学影像(如CT、MRI、病理切片)。它们缺乏处理像素信息、识别图像特征(如病灶形状、纹理、密度)的能力。虽然多模态LLMs(MLLMs)结合了视觉编码器和语言模型,可以在接收图像作为输入后生成文本描述或回答关于图像的问题 6465。但这种处理方式仍是间接的,MLLMs通常是先将影像数据转化为高维特征向量,再由语言模型理解这些特征,而不是直接进行图像分析和模式识别。对于精准的影像诊断和量化分析,仍然需要专门的计算机视觉模型(如CNN)来完成 66。
  3. 缺乏领域特定知识的深度:尽管LLMs通过大量文本数据进行训练,但其对医学领域特定知识的深度理解和推理能力仍不如经过专业领域微调的专家系统或专门的深度学习模型。在某些高度专业化的任务中,LLMs可能会出现“幻觉”或生成不准确的信息。

结合其他模型:
鉴于上述局限性,在多参数建模中,LLMs通常需要与其他机器学习或深度学习模型结合使用,形成“混合AI”策略。例如:

  • LLMs + 传统ML/DL:LLMs负责从非结构化文本中提取和结构化关键信息;这些结构化信息再与临床数值数据、组学数据一同输入到传统机器学习模型(如逻辑回归、随机森林)或深度学习模型(如多层感知机)中进行建模和预测。
  • LLMs + 计算机视觉:LLMs处理文本报告和临床笔记,而计算机视觉模型(如CNN)则处理原始医学影像数据。两个模型可以并行工作,它们的输出(如图像特征、文本提取的结构化信息)在后续的融合层进行整合,共同驱动最终的预测模型。例如,在心血管疾病风险预测中,LLM被用于分析非结构化数据,但图像识别和风险标记的自动化检测仍依赖于其他AI模型 67。

因此,LLMs在多参数建模中扮演着重要角色,尤其是在非结构化文本信息的转化和利用方面,但它们并非万能。它们是强大工具箱中的一部分,需要与其他专业工具协同作用,才能构建出真正全面、高效的医学预测模型。

4. 多参数预测模型的验证与临床转化挑战

4.1 模型验证的关键指标与方法

构建一个有效的多参数预测模型只是第一步,其在临床实践中真正发挥作用,离不开严格的模型验证。验证旨在评估模型预测结果的准确性、可靠性和泛化能力,确保其在未见过的新数据上仍能保持良好性能。对于涉及233例患者的“多参数建模预测晚期肺癌免疫治疗疗效的研究”,模型验证尤为关键,需要遵循一套系统的评估标准和方法。

模型验证的关键指标主要包括区分度、校准度和临床实用性:

  1. 区分度 (Discrimination):
    区分度衡量模型将不同结果(例如,治疗有效与无效)的患者区分开来的能力。最常用的指标是受试者操作特征曲线下面积 (Area Under the Receiver Operating Characteristic Curve, AUC)。

    • AUC值:AUC值范围从0.5(随机猜测)到1.0(完美区分)。一个优秀的预测模型通常要求AUC值在0.75以上,越接近1.0表示区分能力越强。在多参数模型中,通常期望通过融合多维度信息,使AUC显著高于单一标志物模型 2868。例如,在预测非小细胞肺癌新辅助免疫化疗后的病理完全缓解(pCR)时,Delta-RF模型与联合模型在训练集和验证集上的AUC分别为0.74和0.788,0.718和0.737 28;另一项研究中,肿瘤内部异质性模型在训练集和外部验证集的AUC分别为0.861和0.781,均优于传统影像组学模型 69。这些都表明了区分度是衡量模型预测能力的首要指标。
  2. 校准度 (Calibration):
    校准度评估模型预测的概率与实际观察到的事件发生率之间的一致性,即模型预测“概率为X”的事件,实际发生率是否接近X。

    • 校准曲线 (Calibration Curve):通过绘制校准曲线来直观评估模型的校准度。在该曲线上,x轴代表模型预测的事件概率,y轴代表实际观察到的事件发生率。如果模型校准良好,曲线应紧密贴近对角线(y=x)。校准曲线下方的曲线表示模型倾向于高估风险,而上方的曲线则表示低估风险。例如,在预测非小细胞肺癌新辅助免疫化疗后的病理完全缓解时,研究人员通过校准曲线展示了模型预测值与实际观察值之间良好的一致性 28。良好的校准度对于临床决策至关重要,因为它确保了医生可以信任模型给出的概率估计,进而为患者提供准确的风险评估 70。
  3. 临床实用性 (Clinical Utility):
    临床实用性评估模型在临床决策中的实际价值和潜在影响。仅仅有高区分度和校准度不足以说明模型具有临床价值,还需要考虑其是否能指导有效的临床干预。

    • 决策曲线分析 (Decision Curve Analysis, DCA):DCA是一种评估预测模型临床净效益的方法。它通过计算在不同风险阈值下,使用模型所带来的净效益来评估其临床价值,并与“所有患者都治疗”或“所有患者都不治疗”的策略进行比较。决策曲线高于X轴且高于其他简单策略,则表明模型具有临床实用性 2870。一项预测头颈部鳞状细胞癌新辅助化疗免疫治疗后病理完全缓解的模型,其决策曲线分析显示出较高的临床实用性 71。
    • 净重分类指数 (Net Reclassification Index, NRI) 和 整合判别改善 (Integrated Discrimination Improvement, IDI):这些指标用于评估新加入的生物标志物或预测因子对现有模型分类能力的改善程度。

内部验证与外部验证的实施要点:

对于233例患者的数据集,进行充分的内部验证和外部验证至关重要,以评估模型的泛化能力并防止过拟合 7273。

  1. 内部验证 (Internal Validation):
    内部验证是在用于模型开发的数据集内部进行。对于233例患者的数据,可以采用以下方法:

    • 训练集/验证集划分:将数据集随机划分为训练集(约70-80%的数据,如163-186例患者)和内部验证集(约20-30%的数据,如47-70例患者)。模型在训练集上学习参数,在内部验证集上评估性能。这种方法能初步评估模型对新数据的预测能力 2874。
    • 交叉验证 (Cross-Validation):例如,K折交叉验证(K-fold cross-validation),将数据分成K个子集,每次用K-1个子集训练模型,用剩余的1个子集进行验证,重复K次。所有K次验证结果的平均值作为模型性能的估计。这种方法能更充分地利用有限的数据,减少随机划分带来的偏差。对于233例患者,可以考虑进行5折或10折交叉验证 74。
    • Bootstrap重采样 (Bootstrap Resampling):从原始数据集中有放回地随机抽取与原始数据集大小相同的样本,重复多次(例如200-1000次)构建多个训练集,并在未被抽到的样本上进行验证。这能提供更稳健的性能估计和置信区间。
  2. 外部验证 (External Validation):
    外部验证是在独立于模型开发和内部验证的数据集上进行,这是评估模型泛化能力和临床实用性的“金标准” 7273。对于233例患者的数据集,如果条件允许,应尽可能寻求来自不同中心、不同时间段或不同人群的独立数据集进行外部验证。

    • 重要性:外部验证能够发现模型在不同临床环境、数据采集协议或患者特征下可能出现的问题。例如,一项预测黑色素瘤免疫治疗响应和预后的机器学习模型系统综述指出,缺乏外部验证是当前研究的主要局限性之一 75。
    • 实施要点:外部验证数据集应尽可能与训练集具有相似的临床特征和数据质量,但又不能完全相同。如果无法获得独立的外部数据集,可以考虑从233例患者中预留一部分作为独立的外部验证集,但需注意其代表性。例如,有研究将来自不同医院的患者数据作为外部验证集,以评估模型的泛化能力 23697176。

通过上述严格的验证流程,可以全面评估多参数预测模型的性能,确保其在应用于实际临床决策时具备足够的科学依据和可靠性。特别是考虑到“233例患者”的数据规模,合理的内部与外部验证策略将是确保研究结果稳健性的关键。

4.2 临床转化的主要障碍与对策

多参数预测模型从研究阶段迈向临床实际应用的过程中,面临着诸多挑战。尽管模型在验证阶段展现出良好性能,但其在真实世界中的可操作性、可信赖性及广泛采纳性仍需解决。对于233例晚期肺癌免疫治疗患者的多参数模型,尤其需要关注以下几个主要障碍及其应对策略:

  1. 数据标准化与异质性管理:

    • 障碍:临床数据来源于不同的医疗机构,在数据采集协议、设备型号、影像参数、实验室检测方法等方面存在差异。例如,不同医院的CT扫描参数、切片厚度、造影剂使用等可能不同,导致影像数据在纹理、密度上存在差异 77。组学数据(如基因测序)也可能因测序平台、文库制备方法而产生批次效应。这种异质性使得模型在训练数据上表现良好,但在不同来源的外部数据上泛化能力下降,即所谓的“域适应”问题。此外,历史数据的质量不一,包括数据缺失、记录不规范等问题,也给模型训练带来困难。
    • 对策:
      • 统一数据标准与协议:推广采用国际通用的数据标准(如OMOP CDM 78),制定严格的数据采集和处理协议,确保多中心数据的统一性。
      • 数据预处理与归一化:开发先进的数据预处理技术,如影像数据的配准、强度归一化、批次效应校正算法等,以减少数据异质性对模型性能的影响。
      • 联邦学习与隐私保护计算:利用联邦学习(Federated Learning)等技术,允许不同医疗机构在不共享原始患者数据的情况下,共同训练模型。模型参数在各中心本地训练后上传至中央服务器进行聚合,有效保护了患者隐私并克服了数据共享障碍 7879。例如,POPCORN系统就是一个基于多变量元分析和贝叶斯框架的联邦学习平台,它能够在不共享患者层面数据的情况下,构建多中心预测模型,并已在结直肠癌预后预测中得到验证 79。
  2. 模型可解释性与临床接受度:

    • 障碍:许多高性能的多参数模型,特别是深度学习模型,通常被认为是“黑箱模型” 80。其内部决策机制不透明,医生难以理解模型给出预测结果的依据。这种缺乏可解释性使得临床医生难以信任和采纳模型建议,尤其在关乎患者生命的医疗决策中。在实践中,医生可能更倾向于结合自身经验进行直觉判断,而非完全依赖模型输出 81。
    • 对策:
      • 可解释人工智能(XAI)技术:发展和应用XAI方法,如LIME、SHAP、Grad-CAM等,以揭示模型预测结果背后的关键特征和决策路径。例如,通过可视化热图展示影像模型关注的区域,或量化各临床参数对预测结果的贡献度。
      • 透明化设计:在模型设计阶段就考虑可解释性,例如优先选择逻辑回归、决策树等本身具有较好可解释性的模型,或构建混合模型,将“黑箱”模型的预测结果与可解释模型相结合。
      • 人机协作模式:将预测模型作为临床辅助决策工具,而非完全替代人类决策。模型提供风险评估和建议,最终决策权仍由经验丰富的医生掌握,通过人机交互界面清晰呈现模型输出和解释,促进医生对模型的理解和信任 81。
  3. 多中心数据共享与伦理法律框架:

    • 障碍:构建高质量的多参数模型往往需要大规模、多中心的数据集。然而,跨机构数据共享受到严格的隐私保护法规(如GDPR、HIPAA)限制和伦理审查壁垒 82。数据所有权、数据使用协议、以及在数据共享过程中如何确保患者隐私和数据安全,是目前面临的巨大挑战 7883。
    • 对策:
      • 完善法规与政策:推动制定明确、统一的医疗数据共享法律法规和伦理指导原则,为多中心研究提供清晰的合规路径。
      • 隐私保护技术:除了联邦学习,还可以采用差分隐私(Differential Privacy)、同态加密(Homomorphic Encryption)等技术,在数据处理和分析过程中提供强大的隐私保护,降低数据泄露风险。
      • 建立数据共享平台与联盟:鼓励医疗机构、研究中心和科技公司合作,共同建立安全、合规的医疗数据共享平台和研究联盟,在确保患者隐私的前提下,实现数据的有效利用。例如,Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM)就是一种促进多中心数据标准化的常见数据模型,dsOMOP等工具进一步将其与联邦分析平台集成,允许在不转移个人数据的情况下进行安全分析 78。
  4. 模型维护与持续更新:

    • 障碍:疾病特征、治疗方案和临床实践指南都在不断演变,使得预测模型可能会随着时间推移而“过时”,导致性能下降(模型漂移)。例如,新的免疫治疗药物或联合疗法出现,可能改变了疾病的治疗响应模式。
    • 对策:
      • 定期再训练与校准:建立模型性能监测机制,定期对模型进行再训练和校准,利用新的临床数据保持其预测的准确性和时效性。
      • 动态更新机制:开发能够自动学习和适应数据分布变化的动态预测模型,例如基于强化学习或在线学习的模型,使其能够持续优化。
      • 版本管理与迭代:对模型的不同版本进行严格管理,确保临床使用的是最新且经过验证的版本。

解决这些临床转化障碍需要跨学科的协作,包括医学、计算机科学、生物统计学、伦理学和法律领域的专家共同努力,以确保多参数预测模型能够安全、有效地服务于临床实践,最终惠及广大晚期肺癌患者。

内容由 AI 生成,仅供参考,请仔细甄别

参考文献

1First-Line Immunotherapy for Non-Small-Cell Lung Cancer.PubMed

Martin Reck, Jordi Remon, Matthew D Hellmann
J Clin Oncol. 2022 Feb 20;40(6):586-597. doi: 10.1200/JCO.21.01497. Epub 2022 Jan 5.
For patients with metastatic non-small-cell lung cancer (mNSCLC), the last decade has been characterized by critical progress that has contributed to substantially improved survival. In particular, the development of specific antibodies against the programmed death (PD-1) receptor, programmed death-ligand 1 (PD-L1), and the cytotoxic T-lymphocyte-associated protein 4 receptor in the therapeutic strategy of mNSCLC either in first- or in second-line settings have led to unprecedented prolonged survival for a proportion of these patients. Although clinical development of immune checkpoint inhibitors with anti-PD-1 and PD-L1 therapies largely began as monotherapy in the second-line setting, the more recent progress has shifted toward combination approaches in first-line settings as well as the integration of immunotherapy into the clinical paradigm in earlier stages. Today, with the exception of mNSCLC harboring targetable oncogenes, nearly all patients with mNSCLC receive PD-1 or PD-L1 therapy in first-line settings. Here we report the current status of first-line immunotherapy in mNSCLC together with current challenges in selecting the best immunotherapeutic approach for the individual patient.

2Emerging immunotherapy targets in lung cancer.PubMed

Hao-Hua Zhu, Yu Feng, Xing-Sheng Hu
Chin Med J (Engl). 2020 Oct 20;133(20):2456-2465. doi: 10.1097/CM9.0000000000001082.
Immunotherapy has become the mainstay for lung cancer treatment, providing sustained therapeutic responses and improved prognosis compared with those obtained with surgery, chemotherapy, radiotherapy, and targeted therapy. It has the potential for anti-tumor treatment and killing tumor cells by activating human immunity and has moved the targets of anti-cancer therapy from malignant tumor cells to immune cell subsets. Two kinds of immune checkpoints, cytotoxic T-lymphocyte-associated antigen 4 (CTLA-4) and programmed death-1 (PD-1)/programmed death ligand 1 (PD-L1), are the main targets of current immunotherapy in lung cancer. Despite the successful outcomes achieved by immune checkpoint inhibitors, a small portion of lung cancer patients remain unresponsive to checkpoint immunotherapy or may ultimately become resistant to these agents as a result of the complex immune modulatory network in the tumor microenvironment. Therefore, it is imperative to exploit novel immunotherapy targets to further expand the proportion of patients benefiting from immunotherapy. This review summarizes the molecular features, biological function, and clinical significance of several novel checkpoints that have important roles in lung cancer immune responses beyond the CTLA-4 and PD-1/PD-L1 axes, including the markers of co-inhibitory and co-stimulatory T lymphocyte pathways and inhibitory markers of macrophages and natural killer cells.

3Managing Resistance to Immune Checkpoint Inhibitors in Lung Cancer: Treatment and Novel Strategies.PubMed

Antonio Passaro, Julie Brahmer, Scott Antonia, et al.
J Clin Oncol. 2022 Feb 20;40(6):598-610. doi: 10.1200/JCO.21.01845. Epub 2022 Jan 5.
A proportion of patients with lung cancer experience long-term clinical benefit with immune checkpoint inhibitors (ICIs). However, most patients develop disease progression during treatment or after treatment discontinuation. Definitions of immune resistance are heterogeneous according to different clinical and biologic features. Primary resistance and acquired resistance, related to tumor-intrinsic and tumor-extrinsic mechanisms, are identified according to previous response patterns and timing of occurrence. The clinical resistance patterns determine differential clinical approaches. To date, several combination therapies are under development to delay or prevent the occurrence of resistance to ICIs, including the blockade of immune coinhibitory signals, the activation of those with costimulatory functions, the modulation of the tumor microenvironment, and the targeting T-cell priming. Tailoring the specific treatments with distinctive biologic resistance mechanisms would be ideal to improve the design and results of clinical trial. In this review, we reviewed the available evidence on immune resistance mechanisms, clinical definitions, and management of resistance to ICIs in lung cancer. We also reviewed data on novel strategies under investigation in this setting.

4PD-L1 expression as a predictive biomarker in advanced non-small-cell lung cancer: updated survival data.PubMed

Pedro N Aguiar, Ramon Andrade De Mello, Peter Hall, et al.
Immunotherapy. 2017 May;9(6):499-506. doi: 10.2217/imt-2016-0150.
AIM: The treatment of non-small-cell lung cancer has changed after the development of the immune checkpoint inhibitors. Although the most studied biomarker is PD-L1 expression, its clinical significance is still debatable. In this article, we show the updated survival analysis of all published data. METHODS: We searched in network and conference data sources for relevant clinical studies of immunotherapy for non-small-cell lung cancer that assessed the PD-L1 expression even as an exploratory analysis. The updated survival hazard ratios (HR) were included in the analysis. RESULTS: 14 studies with 2857 patients were included (2019 treated with immunotherapy). The response rate was as higher among PD-L1-positive patients (RR: 2.19, 95% CI: 1.63-2.94). PD-L1 expression was also related to better progression-free survival (HR: 0.69, 95% CI: 0.57-0.85) and better overall survival (HR: 0.77, 95% CI: 0.67-0.89). CONCLUSION: PD-L1 overexpression predicts activity as well as better survival for patients treated with immune checkpoint inhibitors.

5Neoadjuvant Chemoimmunotherapy for NSCLC: A Systematic Review and Meta-Analysis.PubMed

Mark Sorin, Connor Prosty, Louis Ghaleb, et al.
JAMA Oncol. 2024 May 1;10(5):621-633. doi: 10.1001/jamaoncol.2024.0057.
IMPORTANCE: To date, no meta-analyses have comprehensively assessed the association of neoadjuvant chemoimmunotherapy with clinical outcomes in non-small cell lung cancer (NSCLC) in randomized and nonrandomized settings. In addition, there exists controversy concerning the efficacy of neoadjuvant chemoimmunotherapy for patients with NSCLC with programmed cell death 1 ligand 1 (PD-L1) levels less than 1%. OBJECTIVE: To compare neoadjuvant chemoimmunotherapy with chemotherapy by adverse events and surgical, pathological, and efficacy outcomes using recently published randomized clinical trials and nonrandomized trials. DATA SOURCES: MEDLINE and Embase were systematically searched from January 1, 2013, to October 25, 2023, for all clinical trials of neoadjuvant chemoimmunotherapy and chemotherapy that included at least 10 patients. STUDY SELECTION: Observational studies and trials reporting the use of neoadjuvant radiotherapy, including chemoradiotherapy, molecular targeted therapy, or immunotherapy monotherapy, were excluded. MAIN OUTCOMES AND MEASURES: Surgical, pathological, and efficacy end points and adverse events were pooled using a random-effects meta-analysis. RESULTS: Among 43 eligible trials comprising 5431 patients (4020 males [74.0%]; median age range, 55-70 years), there were 8 randomized clinical trials with 3387 patients. For randomized clinical trials, pooled overall survival (hazard ratio, 0.65; 95% CI, 0.54-0.79; I2 = 0%), event-free survival (hazard ratio, 0.59; 95% CI, 0.52-0.67; I2 = 14.9%), major pathological response (risk ratio, 3.42; 95% CI, 2.83-4.15; I2 = 31.2%), and complete pathological response (risk ratio, 5.52; 95% CI, 4.25-7.15; I2 = 27.4%) favored neoadjuvant chemoimmunotherapy over neoadjuvant chemotherapy. For patients with baseline tumor PD-L1 levels less than 1%, there was a significant benefit in event-free survival for neoadjuvant chemoimmunotherapy compared with chemotherapy (hazard ratio, 0.74; 95% CI, 0.62-0.89; I2 = 0%). CONCLUSION AND RELEVANCE: This study found that neoadjuvant chemoimmunotherapy was superior to neoadjuvant chemotherapy across surgical, pathological, and efficacy outcomes. These findings suggest that patients with resectable NSCLC with tumor PD-L1 levels less than 1% may have an event-free survival benefit with neoadjuvant chemoimmunotherapy.

6Artificial intelligence-based prediction of clinical outcome in immunotherapy and targeted therapy of lung cancer.PubMed

Xiaomeng Yin, Hu Liao, Hong Yun, et al.
Semin Cancer Biol. 2022 Nov;86(Pt 2):146-159. doi: 10.1016/j.semcancer.2022.08.002. Epub 2022 Aug 11.
Lung cancer accounts for the main proportion of malignancy-related deaths and most patients are diagnosed at an advanced stage. Immunotherapy and targeted therapy have great advances in application in clinics to treat lung cancer patients, yet the efficacy is unstable. The response rate of these therapies varies among patients. Some biomarkers have been proposed to predict the outcomes of immunotherapy and targeted therapy, including programmed cell death-ligand 1 (PD-L1) expression and oncogene mutations. Nevertheless, the detection tests are invasive, time-consuming, and have high demands on tumor tissue. The predictive performance of conventional biomarkers is also unsatisfactory. Therefore, novel biomarkers are needed to effectively predict the outcomes of immunotherapy and targeted therapy. The application of artificial intelligence (AI) can be a possible solution, as it has several advantages. AI can help identify features that are unable to be used by humans and perform repetitive tasks. By combining AI methods with radiomics, pathology, genomics, transcriptomics, proteomics, and clinical data, the integrated model has shown predictive value in immunotherapy and targeted therapy, which significantly improves the precision treatment of lung cancer patients. Herein, we reviewed the application of AI in predicting the outcomes of immunotherapy and targeted therapy in lung cancer patients, and discussed the challenges and future directions in this field.

7Personalized prediction of immunotherapy response in lung cancer patients using advanced radiomics and deep learning.PubMed

Chien-Yi Liao, Yuh-Min Chen, Yu-Te Wu, et al.
Cancer Imaging. 2024 Sep 30;24(1):129. doi: 10.1186/s40644-024-00779-4.
BACKGROUND: Lung cancer (LC) is a leading cause of cancer-related mortality, and immunotherapy (IO) has shown promise in treating advanced-stage LC. However, identifying patients likely to benefit from IO and monitoring treatment response remains challenging. This study aims to develop a predictive model for progression-free survival (PFS) in LC patients with IO based on clinical features and advanced imaging biomarkers. MATERIALS AND METHODS: A retrospective analysis was conducted on a cohort of 206 LC patients receiving IO treatment. Pre-treatment computed tomography images were used to extract advanced imaging biomarkers, including intratumoral and peritumoral-vasculature radiomics. Clinical features, including age, gene status, hematology, and staging, were also collected. Key radiomic and clinical features for predicting IO outcomes were identified using a two-step feature selection process, including univariate Cox regression and chi-squared test, followed by sequential forward selection. The DeepSurv model was constructed to predict PFS based on clinical and radiomic features. Model performance was evaluated using the area under the time-dependent receiver operating characteristic curve (AUC) and concordance index (C-index). RESULTS: Combining radiomics of intratumoral heterogeneity and peritumoral-vasculature with clinical features demonstrated a significant enhancement (p < 0.001) in predicting IO response. The proposed DeepSurv model exhibited a prediction performance with AUCs ranging from 0.76 to 0.80 and a C-index of 0.83. Furthermore, the predicted personalized PFS curves revealed a significant difference (p < 0.05) between patients with favorable and unfavorable prognoses. CONCLUSIONS: Integrating intratumoral and peritumoral-vasculature radiomics with clinical features enabled the development of a predictive model for PFS in LC patients with IO. The proposed model's capability to estimate individualized PFS probability and differentiate the prognosis status held promise to facilitate personalized medicine and improve patient outcomes in LC.

8Predicting immunotherapy response of advanced bladder cancer through a meta-analysis of six independent cohorts.PubMed

Lilian Marie Boll, Sergio Vázquez Montes de Oca, Marta E Camarena, et al.
Nat Commun. 2025 Feb 20;16(1):1213. doi: 10.1038/s41467-025-56462-0.
Advanced bladder cancer patients show very variable responses to immune checkpoint inhibitors (ICIs) and effective strategies to predict response are still lacking. Here we integrate mutation and gene expression data from 707 advanced bladder cancer patients treated with anti-PD-1/anti-PD-L1 to build highly accurate predictive models. We find that, in addition to tumor mutational burden (TMB), enrichment in the APOBEC mutational signature, and the abundance of pro-inflammatory macrophages, are major factors associated with the response. Paradoxically, patients with high immune infiltration do not show an overall better response. We show that this can be explained by the activation of immune suppressive mechanisms in a large portion of these patients. In the case of non-immune-infiltrated cancer subtypes, we uncover specific variables likely to be involved in the response. Our findings provide information for advancing precision medicine in patients with advanced bladder cancer treated with immunotherapy.

9Assessing treatment outcomes of chemoimmunotherapy in extensive-stage small cell lung cancer: an integrated clinical and radiomics approach.PubMed

Jie Zhao, Yayi He, Xue Yang, et al.
J Immunother Cancer. 2023 Sep;11(9). doi: 10.1136/jitc-2023-007492.
BACKGROUND: Small cell lung cancer (SCLC) is a highly malignant cancer characterized by metastasis and an extremely poor prognosis. Although combined chemoimmunotherapy improves the prognosis of extensive-stage (ES)-SCLC, the survival benefits remain limited. Furthermore, no reliable biomarker is available so far to predict the treatment outcomes for chemoimmunotherapy. METHODS: This retrospective study included patients with ES-SCLC treated with first-line combined atezolizumab or durvalumab with standard chemotherapy between Janauray 1, 2019 and October 1, 2022 at five medical centers in China as the chemoimmunotherapy group. The patients were divided into one training cohort and two independent external validation cohorts. Additionally, we created a control group of ES-SCLC who was treated with first-line standard chemotherapy alone. The Radiomics Score was derived using machine learning algorithms based on the radiomics features extracted in the regions of interest delineated on the chest CT obtained before treatment. Cox proportional hazards regression analysis was performed to identify clinical features associated with therapeutic efficacy. The log-rank test, time-dependent receiver operating characteristic curve, and Concordance Index (C-index) were used to assess the effectiveness of the models. RESULTS: A total of 341 patients (mean age, 62±8.7 years) were included in our study. After a median follow-up time of 12.1 months, the median progression-free survival (mPFS) was 7.1 (95% CI 6.6 to 7.7) months, whereas the median overall survival (mOS) was not reached. The TNM stage, Eastern Cooperative Oncology Group performance status, and Lung Immune Prognostic Index showed significant correlations with PFS. We proposed a predictive model based on eight radiomics features to determine the risk of chemoimmunotherapy resistance among patients with SCLC (validation set 1: mPFS, 12.0 m vs 5.0 m, C-index=0.634; validation set 2: mPFS, 10.8 m vs 6.1 m, C-index=0.617). By incorporating the clinical features associated with PFS into the radiomics model, the predictive efficacy was substantially improved. Consequently, the low-progression-risk group exhibited a significantly longer mPFS than the high-progression-risk group in both validation set 1 (mPFS, 12.8 m vs 4.5 m, HR=0.40, p=0.028) and validation set 2 (mPFS, 9.2 m vs 4.6 m, HR=0.30, p=0.012). External validation set 1 and set 2 yielded the highest 6-month area under the curve and C-index of 0.852 and 0.820, respectively. Importantly, the integrated prediction model also exhibited considerable differentiation power for survival outcomes. The HR for OS derived from the low-progression-risk and high-progression-risk groups was 0.28 (95% CI 0.17 to 0.48) in all patients and 0.20 (95% CI 0.08 to 0.54) in validation set. By contrast, no significant differences were observed in PFS and OS, between high-progression-risk patients receiving chemoimmunotherapy and the chemotherapy cohort (mPFS, 5.5 m vs 5.9 m, HR=0.90, p=0.547; mOS, 14.5 m vs 13.7 m, HR=0.97, p=0.910). CONCLUSIONS: The integrated clinical and radiomics model can predict the treatment outcomes in patients with ES-SCLC receiving chemoimmunotherapy, rendering a convenient and low-cost prognostic model for decision-making regarding patient management.

10Predicting benefit from immune checkpoint inhibitors in patients with non-small-cell lung cancer by CT-based ensemble deep learning: a retrospective study.PubMed

Maliazurina B Saad, Lingzhi Hong, Muhammad Aminu, et al.
Lancet Digit Health. 2023 Jul;5(7):e404-e420. doi: 10.1016/S2589-7500(23)00082-1. Epub 2023 May 31.
BACKGROUND: Only around 20-30% of patients with non-small-cell lung cancer (NCSLC) have durable benefit from immune-checkpoint inhibitors. Although tissue-based biomarkers (eg, PD-L1) are limited by suboptimal performance, tissue availability, and tumour heterogeneity, radiographic images might holistically capture the underlying cancer biology. We aimed to investigate the application of deep learning on chest CT scans to derive an imaging signature of response to immune checkpoint inhibitors and evaluate its added value in the clinical context. METHODS: In this retrospective modelling study, 976 patients with metastatic, EGFR/ALK negative NSCLC treated with immune checkpoint inhibitors at MD Anderson and Stanford were enrolled from Jan 1, 2014, to Feb 29, 2020. We built and tested an ensemble deep learning model on pretreatment CTs (Deep-CT) to predict overall survival and progression-free survival after treatment with immune checkpoint inhibitors. We also evaluated the added predictive value of the Deep-CT model in the context of existing clinicopathological and radiological metrics. FINDINGS: Our Deep-CT model demonstrated robust stratification of patient survival of the MD Anderson testing set, which was validated in the external Stanford set. The performance of the Deep-CT model remained significant on subgroup analyses stratified by PD-L1, histology, age, sex, and race. In univariate analysis, Deep-CT outperformed the conventional risk factors, including histology, smoking status, and PD-L1 expression, and remained an independent predictor after multivariate adjustment. Integrating the Deep-CT model with conventional risk factors demonstrated significantly improved prediction performance, with overall survival C-index increases from 0·70 (clinical model) to 0·75 (composite model) during testing. On the other hand, the deep learning risk scores correlated with some radiomics features, but radiomics alone could not reach the performance level of deep learning, indicating that the deep learning model effectively captured additional imaging patterns beyond known radiomics features. INTERPRETATION: This proof-of-concept study shows that automated profiling of radiographic scans through deep learning can provide orthogonal information independent of existing clinicopathological biomarkers, bringing the goal of precision immunotherapy for patients with NSCLC closer. FUNDING: National Institutes of Health, Mark Foundation Damon Runyon Foundation Physician Scientist Award, MD Anderson Strategic Initiative Development Program, MD Anderson Lung Moon Shot Program, Andrea Mugnaini, and Edward L C Smith.

11Prediction of Response to FDA-Approved Targeted Therapy and Immunotherapy in Acute Lymphoblastic Leukemia (ALL).PubMed

Zakaria Yahya Khawaji, Nussaiba Yahya Khawaji, Mohammed Abdullah Alahmadi, et al.
Curr Treat Options Oncol. 2024 Sep;25(9):1163-1183. doi: 10.1007/s11864-024-01237-w. Epub 2024 Aug 5.
Acute lymphoblastic leukemia (ALL) represents the predominant cancer in pediatric populations, though its occurrence in adults is relatively rare. Pre-treatment risk stratification is crucial for predicting prognosis. Important factors for assessment include patient age, white blood cell (WBC) count at diagnosis, extramedullary involvement, immunophenotype, and cytogenetic aberrations. Minimal residual disease (MRD), primarily assessed by flow cytometry following remission, plays a substantial role in guiding management plans. Over the past decade, significant advancements in ALL outcomes have been witnessed. Conventional chemotherapy has remarkably reduced mortality rates; however, its intensive nature raises safety concerns and has led to the emergence of treatment-resistant cases with recurrence of relapses. Consequently, The U.S. Food and Drug Administration (FDA) has approved several novel treatments for relapsed/refractory ALL due to their demonstrated efficacy, as indicated by improved complete remission and survival rates. These treatments include tyrosine kinase inhibitors (TKIs), the anti-CD19 monoclonal antibody blinatumomab, anti-CD22 inotuzumab ozogamicin, anti-CD20 rituximab, and chimeric antigen receptor (CAR) T-cell therapy. Identifying the variables that influence treatment decisions is a pressing necessity for tailoring therapy based on heterogeneous patient characteristics. Key predictive factors identified in various observational studies and clinical trials include prelymphodepletion disease burden, complex genetic abnormalities, and MRD. Furthermore, the development of serious adverse events following treatment could be anticipated through predictive models, allowing for appropriate prophylactic measures to be considered. The ultimate aim is to incorporate the concept of precision medicine in the field of ALL through valid prediction platform to facilitate the selection of the most suitable treatment approach.

12Microsatellite Instability, Mismatch Repair, and Tumor Mutation Burden in Lung Cancer.PubMed

Oana C Rosca, Oana E Vele
Surg Pathol Clin. 2024 Jun;17(2):295-305. doi: 10.1016/j.path.2023.11.011. Epub 2023 Dec 20.
Since US Food and Drug Administration approval of programmed death ligand 1 (PD-L1) as the first companion diagnostic for immune checkpoint inhibitors (ICIs) in non-small cell lung cancer, many patients have experienced increased overall survival. To improve selection of ICI responders versus nonresponders, microsatellite instability/mismatch repair deficiency (MSI/MMR) and tumor mutation burden (TMB) came into play. Clinical data show PD-L1, MSI/MMR, and TMB are independent predictive immunotherapy biomarkers. Harmonization of testing methodologies, optimization of assay design, and results analysis are ongoing. Future algorithms to determine immunotherapy eligibility might involve complementary use of current and novel biomarkers. Artificial intelligence could facilitate algorithm implementation to convert complex genetic data into recommendations for specific ICIs.

13Predictive Biomarkers for Immunotherapy in Lung Cancer: Perspective From the International Association for the Study of Lung Cancer Pathology Committee.PubMed

Mari Mino-Kenudson, Kurt Schalper, Wendy Cooper, et al.
J Thorac Oncol. 2022 Dec;17(12):1335-1354. doi: 10.1016/j.jtho.2022.09.109. Epub 2022 Sep 29.
Immunotherapy including immune checkpoint inhibitors (ICIs) has become the backbone of treatment for most lung cancers with advanced or metastatic disease. In addition, they have increasingly been used for early stage tumors in neoadjuvant and adjuvant settings. Unfortunately, however, only a subset of patients experiences meaningful response to ICIs. Although programmed death-ligand 1 (PD-L1) protein expression by immunohistochemistry (IHC) has played a role as the principal predictive biomarker for immunotherapy, its performance may not be optimal, and it suffers multiple practical issues with different companion diagnostic assays approved. Similarly, tumor mutational burden (TMB) has multiple technical issues as a predictive biomarker for ICIs. Now, ongoing research on tumor- and host immune-specific factors has identified immunotherapy biomarkers that may provide better response and prognosis prediction, in particular in a multimodal approach. This review by the International Association for the Study of Lung Cancer Pathology Committee provides an overview of various immunotherapy biomarkers, including updated data on PD-L1 IHC and TMB, and assessments of neoantigens, genetic and epigenetic signatures, immune microenvironment by IHC and transcriptomics, and microbiome and pathologic response to neoadjuvant immunotherapies. The aim of this review is to underline the efficacy of new individual or combined predictive biomarkers beyond PD-L1 IHC and TMB.

14Oncogene-specific differences in tumor mutational burden, PD-L1 expression, and outcomes from immunotherapy in non-small cell lung cancer.PubMed

Marcelo V Negrao, Ferdinandos Skoulidis, Meagan Montesion, et al.
J Immunother Cancer. 2021 Aug;9(8). doi: 10.1136/jitc-2021-002891.
BACKGROUND: Non-small cell lung cancer (NSCLC) patients bearing targetable oncogene alterations typically derive limited benefit from immune checkpoint blockade (ICB), which has been attributed to low tumor mutation burden (TMB) and/or PD-L1 levels. We investigated oncogene-specific differences in these markers and clinical outcome. METHODS: Three cohorts of NSCLC patients with oncogene alterations (n=4189 total) were analyzed. Two clinical cohorts of advanced NSCLC patients treated with ICB monotherapy [MD Anderson (MDACC; n=172) and Flatiron Health-Foundation Medicine Clinico-Genomic Database (CGDB; n=894 patients)] were analyzed for clinical outcome. The FMI biomarker cohort (n=4017) was used to assess the association of oncogene alterations with TMB and PD-L1 expression. RESULTS: High PD-L1 expression (PD-L1 ≥50%) rate was 19%-20% in classic , exon 20 and -mutant tumors, and 34%-55% in tumors with , V600E, , , or alterations. Compared with mutant tumors, non-V600E group had higher TMB (9.6 vs 7.8 mutations/Mb, p=0.003), while all other oncogene groups had lower TMB (p<0.001). In the two clinical cohorts treated with ICB, molecular groups with , , , , , or alterations had short progression-free survival (PFS; 1.8-3.7 months), while V600E group was associated with greater clinical benefit from ICB (CGDB cohort: PFS 9.8 months vs 3.7 months, HR 0.66, p=0.099; MDACC cohort: response rate 62% vs 24%; PFS 7.4 vs 2.8 months, HR 0.36, p=0.026). G12C and non-G12C subgroups had similar clinical benefit from ICB in both cohorts. In a multivariable analysis, V600E mutation (HR 0.58, p=0.041), PD-L1 expression (HR 0.57, p=0.022), and high TMB (HR 0.66, p<0.001) were associated with longer PFS. CONCLUSIONS: High TMB and PD-L1 expression are predictive for benefit from ICB treatment in oncogene-driven NSCLCs. NSCLC harboring mutations demonstrated superior benefit from ICB that may be attributed to higher TMB and higher PD-L1 expression in these tumors. Meanwhile and mutations and , , , and fusions define NSCLC subsets with minimal benefit from ICB despite high PD-L1 expression in NSCLC harboring oncogene fusions. These findings indicate a TMB/PD-L1-independent impact on sensitivity to ICB for certain oncogene alterations.

15A Novel Radiogenomics Biomarker for Predicting Treatment Response and Pneumotoxicity From Programmed Cell Death Protein or Ligand-1 Inhibition Immunotherapy in NSCLC.PubMed

Mitchell Chen, Haonan Lu, Susan J Copley, et al.
J Thorac Oncol. 2023 Jun;18(6):718-730. doi: 10.1016/j.jtho.2023.01.089. Epub 2023 Feb 10.
INTRODUCTION: Patient selection for checkpoint inhibitor immunotherapy is currently guided by programmed death-ligand 1 (PD-L1) expression obtained from immunohistochemical staining of tumor tissue samples. This approach is susceptible to limitations resulting from the dynamic and heterogeneous nature of cancer cells and the invasiveness of the tissue sampling procedure. To address these challenges, we developed a novel computed tomography (CT) radiomic-based signature for predicting disease response in patients with NSCLC undergoing programmed cell death protein 1 (PD-1) or PD-L1 checkpoint inhibitor immunotherapy. METHODS: This retrospective study comprises a total of 194 patients with suitable CT scans out of 340. Using the radiomic features computed from segmented tumors on a discovery set of 85 contrast-enhanced chest CTs of patients diagnosed with having NSCLC and their CD274 count, RNA expression of the protein-encoding gene for PD-L1, as the response vector, we developed a composite radiomic signature, lung cancer immunotherapy-radiomics prediction vector (LCI-RPV). This was validated in two independent testing cohorts of 66 and 43 patients with NSCLC treated with PD-1 or PD-L1 inhibition immunotherapy, respectively. RESULTS: LCI-RPV predicted PD-L1 positivity in both NSCLC testing cohorts (area under the curve [AUC] = 0.70, 95% confidence interval [CI]: 0.57-0.84 and AUC = 0.70, 95% CI: 0.46-0.94). In one cohort, it also demonstrated good prediction of cases with high PD-L1 expression exceeding key treatment thresholds (>50%: AUC = 0.72, 95% CI: 0.59-0.85 and >90%: AUC = 0.66, 95% CI: 0.45-0.88), the tumor's objective response to treatment at 3 months (AUC = 0.68, 95% CI: 0.52-0.85), and pneumonitis occurrence (AUC = 0.64, 95% CI: 0.48-0.80). LCI-RPV achieved statistically significant stratification of the patients into a high- and low-risk survival group (hazard ratio = 2.26, 95% CI: 1.21-4.24, p = 0.011 and hazard ratio = 2.45, 95% CI: 1.07-5.65, p = 0.035). CONCLUSIONS: A CT radiomics-based signature developed from response vector CD274 can aid in evaluating patients' suitability for PD-1 or PD-L1 checkpoint inhibitor immunotherapy in NSCLC.

16Neoadjuvant PD-1 Blockade in Resectable Lung Cancer.PubMed

Patrick M Forde, Jamie E Chaft, Kellie N Smith, et al.
N Engl J Med. 2018 May 24;378(21):1976-1986. doi: 10.1056/NEJMoa1716078. Epub 2018 Apr 16.
BACKGROUND: Antibodies that block programmed death 1 (PD-1) protein improve survival in patients with advanced non-small-cell lung cancer (NSCLC) but have not been tested in resectable NSCLC, a condition in which little progress has been made during the past decade. METHODS: In this pilot study, we administered two preoperative doses of PD-1 inhibitor nivolumab in adults with untreated, surgically resectable early (stage I, II, or IIIA) NSCLC. Nivolumab (at a dose of 3 mg per kilogram of body weight) was administered intravenously every 2 weeks, with surgery planned approximately 4 weeks after the first dose. The primary end points of the study were safety and feasibility. We also evaluated the tumor pathological response, expression of programmed death ligand 1 (PD-L1), mutational burden, and mutation-associated, neoantigen-specific T-cell responses. RESULTS: Neoadjuvant nivolumab had an acceptable side-effect profile and was not associated with delays in surgery. Of the 21 tumors that were removed, 20 were completely resected. A major pathological response occurred in 9 of 20 resected tumors (45%). Responses occurred in both PD-L1-positive and PD-L1-negative tumors. There was a significant correlation between the pathological response and the pretreatment tumor mutational burden. The number of T-cell clones that were found in both the tumor and peripheral blood increased systemically after PD-1 blockade in eight of nine patients who were evaluated. Mutation-associated, neoantigen-specific T-cell clones from a primary tumor with a complete response on pathological assessment rapidly expanded in peripheral blood at 2 to 4 weeks after treatment; some of these clones were not detected before the administration of nivolumab. CONCLUSIONS: Neoadjuvant nivolumab was associated with few side effects, did not delay surgery, and induced a major pathological response in 45% of resected tumors. The tumor mutational burden was predictive of the pathological response to PD-1 blockade. Treatment induced expansion of mutation-associated, neoantigen-specific T-cell clones in peripheral blood. (Funded by Cancer Research Institute-Stand Up 2 Cancer and others; ClinicalTrials.gov number, NCT02259621 .).

17Development of tumor mutation burden as an immunotherapy biomarker: utility for the oncology clinic.PubMed

T A Chan, M Yarchoan, E Jaffee, et al.
Ann Oncol. 2019 Jan 1;30(1):44-56. doi: 10.1093/annonc/mdy495.
BACKGROUND: Treatment with immune checkpoint blockade (ICB) with agents such as anti-programmed cell death protein 1 (PD-1), anti-programmed death-ligand 1 (PD-L1), and/or anti-cytotoxic T-lymphocyte-associated protein 4 (CTLA-4) can result in impressive response rates and durable disease remission but only in a subset of patients with cancer. Expression of PD-L1 has demonstrated utility in selecting patients for response to ICB and has proven to be an important biomarker for patient selection. Tumor mutation burden (TMB) is emerging as a potential biomarker. However, refinement of interpretation and contextualization is required. MATERIALS AND METHODS: In this review, we outline the evolution of TMB as a biomarker in oncology, delineate how TMB can be applied in the clinic, discuss current limitations as a diagnostic test, and highlight mechanistic insights unveiled by the study of TMB. We review available data to date studying TMB as a biomarker for response to ICB by tumor type, focusing on studies proposing a threshold for TMB as a predictive biomarker for ICB activity. RESULTS: High TMB consistently selects for benefit with ICB therapy. In lung, bladder and head and neck cancers, the current predictive TMB thresholds proposed approximate 200 non-synonymous somatic mutations by whole exome sequencing (WES). PD-L1 expression influences response to ICB in high TMB tumors with single agent PD-(L)1 antibodies; however, response may not be dependent on PD-L1 expression in the setting of anti-CTLA4 or anti-PD-1/CTLA-4 combination therapy. Disease-specific TMB thresholds for effective prediction of response in various other malignancies are not well established. CONCLUSIONS: TMB, in concert with PD-L1 expression, has been demonstrated to be a useful biomarker for ICB selection across some cancer types; however, further prospective validation studies are required. TMB determination by selected targeted panels has been correlated with WES. Calibration and harmonization will be required for optimal utility and alignment across all platforms currently used internationally. Key challenges will need to be addressed before broader use in different tumor types.

18Genomic Features of Response to Combination Immunotherapy in Patients with Advanced Non-Small-Cell Lung Cancer.PubMed

Matthew D Hellmann, Tavi Nathanson, Hira Rizvi, et al.
Cancer Cell. 2018 May 14;33(5):843-852.e4. doi: 10.1016/j.ccell.2018.03.018. Epub 2018 Apr 12.
Combination immune checkpoint blockade has demonstrated promising benefit in lung cancer, but predictors of response to combination therapy are unknown. Using whole-exome sequencing to examine non-small-cell lung cancer (NSCLC) treated with PD-1 plus CTLA-4 blockade, we found that high tumor mutation burden (TMB) predicted improved objective response, durable benefit, and progression-free survival. TMB was independent of PD-L1 expression and the strongest feature associated with efficacy in multivariable analysis. The low response rate in TMB low NSCLCs demonstrates that combination immunotherapy does not overcome the negative predictive impact of low TMB. This study demonstrates the association between TMB and benefit to combination immunotherapy in NSCLC. TMB should be incorporated in future trials examining PD-(L)1 with CTLA-4 blockade in NSCLC.

19Long-term benefit of immunotherapy in a patient with squamous lung cancer exhibiting mismatch repair deficient/high microsatellite instability/high tumor mutational burden: A case report and literature review.PubMed

Na Li, Zixuan Wan, Dongyan Lu, et al.
Front Immunol. 2023 Jan 10;13:1088683. doi: 10.3389/fimmu.2022.1088683. eCollection 2022.
Genetic mutations that render mismatch repair defective may result in microsatellite instability, which is common in colorectal carcinomas and gastric cancers as well as Lynch syndrome. Mismatch repair deficiency/high microsatellite instability (dMMR/MSI-H) predicts the tumor response to immune checkpoint inhibitors. However, few studies have evaluated the efficacy of immune checkpoint inhibitors in non-small cell lung cancer (NSCLC) patients with dMMR/MSI-H. In this work, we present a patient with advanced squamous lung cancer with dMMR/MSI-H and a high tumor mutational burden (TMB-H) who obtained a long-term benefit from immunotherapy. NSCLC patients with dMMR/MSI-H/TMB-H may thus benefit from immune checkpoint inhibitors.

20Predictive Biomarkers for Immunotherapy in Gastric Cancer: Current Status and Emerging Prospects.PubMed

Wanting Hou, Yaqin Zhao, Hong Zhu
Int J Mol Sci. 2023 Oct 18;24(20):15321. doi: 10.3390/ijms242015321.
Gastric cancer presents substantial management challenges, and the advent of immunotherapy has ignited renewed hope among patients. Nevertheless, a significant proportion of patients do not respond to immunotherapy, and adverse events associated with immunotherapy also occur on occasion, underscoring the imperative to identify suitable candidates for treatment. Several biomarkers, including programmed death ligand-1 expression, tumor mutation burden, mismatch repair status, Epstein-Barr Virus infection, circulating tumor DNA, and tumor-infiltrating lymphocytes, have demonstrated potential in predicting the effectiveness of immunotherapy in gastric cancer. However, the quest for the optimal predictive biomarker for gastric cancer immunotherapy remains challenging, as each biomarker carries its own limitations. Recently, multi-omics technologies have emerged as promising platforms for discovering novel biomarkers that may help in selecting gastric cancer patients likely to respond to immunotherapy. The identification of reliable predictive biomarkers for immunotherapy in gastric cancer holds the promise of enhancing patient selection and improving treatment outcomes. In this review, we aim to provide an overview of clinically established biomarkers of immunotherapy in gastric cancer. Additionally, we introduce newly reported biomarkers based on multi-omics studies in the context of gastric cancer immunotherapy, thereby contributing to the ongoing efforts to refine patient stratification and treatment strategies.

21Comparison of different predictive biomarker testing assays for PD-1/PD-L1 checkpoint inhibitors response: a systematic review and network meta-analysis.PubMed

Haotong Shi, Wenxia Zhang, Lin Zhang, et al.
Front Immunol. 2023 Sep 26;14:1265202. doi: 10.3389/fimmu.2023.1265202. eCollection 2023.
BACKGROUND: Accurate prediction of efficacy of programmed cell death 1 (PD-1)/programmed cell death ligand 1 (PD-L1) checkpoint inhibitors is of critical importance. To address this issue, a network meta-analysis (NMA) comparing existing common measurements for curative effect of PD-1/PD-L1 monotherapy was conducted. METHODS: We searched PubMed, Embase, the Cochrane Library database, and relevant clinical trials to find out studies published before Feb 22, 2023 that use PD-L1 immunohistochemistry (IHC), tumor mutational burden (TMB), gene expression profiling (GEP), microsatellite instability (MSI), multiplex IHC/immunofluorescence (mIHC/IF), other immunohistochemistry and hematoxylin-eosin staining (other IHC&HE) and combined assays to determine objective response rates to anti-PD-1/PD-L1 monotherapy. Study-level data were extracted from the published studies. The primary goal of this study was to evaluate the predictive efficacy and rank these assays mainly by NMA, and the second objective was to compare them in subgroup analyses. Heterogeneity, quality assessment, and result validation were also conducted by meta-analysis. FINDINGS: 144 diagnostic index tests in 49 studies covering 5322 patients were eligible for inclusion. mIHC/IF exhibited highest sensitivity (0.76, 95% CI: 0.57-0.89), the second diagnostic odds ratio (DOR) (5.09, 95% CI: 1.35-13.90), and the second superiority index (2.86). MSI had highest specificity (0.90, 95% CI: 0.85-0.94), and DOR (6.79, 95% CI: 3.48-11.91), especially in gastrointestinal tumors. Subgroup analyses by tumor types found that mIHC/IF, and other IHC&HE demonstrated high predictive efficacy for non-small cell lung cancer (NSCLC), while PD-L1 IHC and MSI were highly efficacious in predicting the effectiveness in gastrointestinal tumors. When PD-L1 IHC was combined with TMB, the sensitivity (0.89, 95% CI: 0.82-0.94) was noticeably improved revealed by meta-analysis in all studies. INTERPRETATION: Considering statistical results of NMA and clinical applicability, mIHC/IF appeared to have superior performance in predicting response to anti PD-1/PD-L1 therapy. Combined assays could further improve the predictive efficacy. Prospective clinical trials involving a wider range of tumor types are needed to establish a definitive gold standard in future.

22Clinical multi-dimensional prognostic nomogram for predicting the efficacy of immunotherapy in NSCLC.PubMed

Qian Zhao, Xiao Zhong, Xiaoqing Wang, et al.
Sci Rep. 2024 Sep 13;14(1):21380. doi: 10.1038/s41598-024-72760-x.
The advent of immunotherapy has greatly improved the prognosis of non-small cell lung (NSCLC) patients. However, given its low response rate and high cost of treatment, the search for valuable predictive markers of treatment efficacy is necessary. Considering the complexity and heterogeneity of the tumour and tumour microenvironment, the construction of a multi-dimensional prediction model is necessary. Therefore, we aimed to integrate clinical parameters, radiomic features, and immune signature data from NSCLC patients receiving immunotherapy to construct a multi-dimensional prediction model to better predict the efficacy of immunotherapy. The current study enrolled 137 NSCLC patients who received immunotherapy. We collected baseline clinical information, CT images, and tumour tissue specimens. Using 3D-Slicer software, radiomic features were extracted from patient CT images, and tumor tissue samples obtained before immunotherapy were subjected to immunohistochemical staining. Then, the least absolute shrinkage and selection operator (LASSO) Cox regression analysis was applied to downscale the data, and the radiomic features and immune signatures associated with the prognosis of immunotherapy patients were identified. The modified lung immune predictive index (mLIPI), radiomics score (Radioscore), immune score and multi-dimensional model nomogram were constructed. The C-index and area under the curve (AUC) were applied to evaluate the predictive efficacy of the models. Three radiomic features and three immune signatures that could predict the efficacy of immunotherapy were eventually screened. Multivariate analysis showed that the mLIPI, Radioscore, and immune score were independent predictive factors for PFS and OS (P < 0.05 for all models). The multi-dimensional model combining the three models showed better predictive efficacy than the mLIPI, Radioscore, and immune score (PFS: 0.721 vs. 0.662 vs. 0.610 vs. 0.610; OS: 0.727 vs. 0.661 vs. 0.601 vs. 0.602 respectively). The multi-dimensional model showed the best predictive efficacy, with C-index for PFS and OS higher than mLIPI, radioscore and immune score: 0.721 vs. 0.662 vs. 0.610 vs. 0.610 for PFS and 0.727 vs. 0.661 vs. 0.601 vs. 0.602 for OS, respectively. The AUC for the multi-dimensional model also performed better than those of the individual models: 0.771 vs. 0.684 vs. 0.715 vs. 0.711 for PFS and 0.768 vs. 0.662 vs. 0.661 vs. 0.658 for OS, respectively. The multi-dimensional model combining the three models had better predictive efficacy than any single model and was more likely to help provide patients personalized and precision medicine.

23Radiomics Biomarkers to Predict Checkpoint Inhibitor Pneumonitis in Non-small Cell Lung Cancer.PubMed

Yonghao Du, Shuo Zhang, Xiaohui Jia, et al.
Acad Radiol. 2025 Mar;32(3):1685-1695. doi: 10.1016/j.acra.2024.09.053. Epub 2024 Oct 11.
RATIONALE AND OBJECTIVES: Immune checkpoint inhibitors (ICIs) have revolutionized the treatment of non-small cell lung cancer (NSCLC). However, immune-related adverse events still occur, of which checkpoint inhibitor pneumonitis (CIP) is the most common. We aimed to construct and validate a contrast-enhanced computed tomography-based radiomic nomogram to predict the probability of CIP before ICIs treatment in NSCLC. MATERIALS AND METHODS: We retrospectively analyzed 685 patients with NSCLC who were initially treated with ICIs. A total of 186 patients were included in our study, and an additional 52 patients from another hospital were considered for external validation. After radiomics feature extraction and selection, we applied a support vector machine classification model to distinguish CIP and used the probability as a radiomics signature. A radiomics-clinical logistic regression model was built using the filtered clinical parameters and a radiomic signature. Receiver operating characteristic, area under the curve (AUC), calibration curve, and decision curve analysis was used for inter-model comparison. RESULTS: The combined radiomics-clinical model constructed using age, interstitial lung disease, emphysema at baseline, and radiomics signature showed an AUC of 0.935, 0.905, and 0.923 for the training, validation, and external validation cohorts, respectively. Compared with the clinical-only (AUC of 0.829, 0.826, and 0.809) and radiomics-only models (0.865, 0.847, and 0.841), the radiomics-clinical displayed better predictive power. CONCLUSION: This combined radiomics-clinical model predicted the probability of CIP during ICIs treatment in patients with NSCLC with favorable accuracy and could therefore be used as an effective tool to guide clinical ICIs decisions.

24Neoadjuvant sintilimab plus chemotherapy in EGFR-mutant NSCLC: Phase 2 trial interim results (NEOTIDE/CTONG2104).PubMed

Chao Zhang, Yu-Xuan Sun, Ding-Cheng Yi, et al.
Cell Rep Med. 2024 Jul 16;5(7):101615. doi: 10.1016/j.xcrm.2024.101615. Epub 2024 Jun 18.
The clinical efficacy of neoadjuvant immunotherapy plus chemotherapy remains elusive in localized epidermal growth factor receptor (EGFR)-mutant non-small cell lung cancer (NSCLC). Here, we report interim results of a Simon's two-stage design, phase 2 trial using neoadjuvant sintilimab with carboplatin and nab-paclitaxel in resectable EGFR-mutant NSCLC. All 18 patients undergo radical surgery, with one patient experiencing surgery delay. Fourteen patients exhibit confirmed radiological response, with 44% achieving major pathological response (MPR) and no pathological complete response (pCR). Similar genomic alterations are observed before and after treatment without influencing the efficacy of subsequent EGFR-tyrosine kinase inhibitors (TKIs) in vitro. Infiltration and T cell receptor (TCR) clonal expansion of CCR8 regulatory T (Treg)/CXCL13 exhausted T (Tex) cells define a subtype of EGFR-mutant NSCLC highly resistant to immunotherapy, with the phenotype potentially serving as a promising signature to predict immunotherapy efficacy. Informed circulating tumor DNA (ctDNA) detection in EGFR-mutant NSCLC could help identify patients nonresponsive to neoadjuvant immunochemotherapy. These findings provide supportive data for the utilization of neoadjuvant immunochemotherapy and insight into immune resistance in EGFR-mutant NSCLC.

25Comprehensively Analyze the Prognosis Significance and Immune Implication of PTPRO in Lung Adenocarcinoma.PubMed

Zhimin Lin, Jinjun Zhang, Biqiong Liu, et al.
Mediators Inflamm. 2023 Feb 9;2023:5248897. doi: 10.1155/2023/5248897. eCollection 2023.
Immunotherapy for lung adenocarcinoma (LUAD) is considered to be a promising treatment option, but only a minority of patients benefit from it. Therefore, it is essential to clarify the regulation mechanism of the tumor immune microenvironment (TIM) of the LUAD. Receptor-type protein tyrosine phosphatase (PTPRO) has been shown to be a tumor suppressor in a variety of tumor; however, its role in LUAD has never been reported. In this study, we first found that PTPRO was lowly expressed in LUAD and positively correlated with patient prognosis. Next, we investigated the relationship between PTPRO and clinical characteristics, and the results showed that gender, age, T, and stage were closely related to the expression level of PTPRO. Moreover, we performed univariate and multivariate analyses, and the results revealed that PTPRO was a protective factor for LUAD. By constructing a nomogram based on the expression level of PTPRO and various clinical characteristics, it was proved that the nomogram has a good predictive capacity. Furthermore, we analyzed the coexpression network of PTPRO through multiple databases and performed GO and KEGG enrichment analyses. The results demonstrated that PTPRO was involved in the regulation of multiple immune pathways. In addition, we analyzed whether PTPRO expression of LUAD regulate immune cell infiltration and the results demonstrated that PTPRO was closely related to the infiltration of various immune cells. Finally, we predicted LUAD sensitivity to chemotherapeutics and response to immunotherapy by PTPRO expression levels. The results showed that PTPRO expression level affect the sensitivity of various chemotherapeutic drugs and may be involved in the efficacy of immunotherapy. These results we obtained suggested that PTPRO is closely related to the prognosis and TIM of LUAD, which may be a potential immunotherapeutic target for LUAD.

26Custom scoring based on ecological topology of gut microbiota associated with cancer immunotherapy outcome.PubMed

Lisa Derosa, Valerio Iebba, Carolina Alves Costa Silva, et al.
Cell. 2024 Jun 20;187(13):3373-3389.e16. doi: 10.1016/j.cell.2024.05.029.
The gut microbiota influences the clinical responses of cancer patients to immunecheckpoint inhibitors (ICIs). However, there is no consensus definition of detrimental dysbiosis. Based on metagenomics (MG) sequencing of 245 non-small cell lung cancer (NSCLC) patient feces, we constructed species-level co-abundance networks that were clustered into species-interacting groups (SIGs) correlating with overall survival. Thirty-seven and forty-five MG species (MGSs) were associated with resistance (SIG1) and response (SIG2) to ICIs, respectively. When combined with the quantification of Akkermansia species, this procedure allowed a person-based calculation of a topological score (TOPOSCORE) that was validated in an additional 254 NSCLC patients and in 216 genitourinary cancer patients. Finally, this TOPOSCORE was translated into a 21-bacterial probe set-based qPCR scoring that was validated in a prospective cohort of NSCLC patients as well as in colorectal and melanoma patients. This approach could represent a dynamic diagnosis tool for intestinal dysbiosis to guide personalized microbiota-centered interventions.

27Combination of computed tomography imaging-based radiomics and clinicopathological characteristics for predicting the clinical benefits of immune checkpoint inhibitors in lung cancer.PubMed

Bin Yang, Li Zhou, Jing Zhong, et al.
Respir Res. 2021 Jun 28;22(1):189. doi: 10.1186/s12931-021-01780-2.
BACKGROUND: In this study, we tested whether a combination of radiomic features extracted from baseline pre-immunotherapy computed tomography (CT) images and clinicopathological characteristics could be used as novel noninvasive biomarkers for predicting the clinical benefits of non-small cell lung cancer (NSCLC) patients treated with immune checkpoint inhibitors (ICIs). METHODS: The data from 92 consecutive patients with lung cancer who had been treated with ICIs were retrospectively analyzed. In total, 88 radiomic features were selected from the pretreatment CT images for the construction of a random forest model. Radiomics model 1 was constructed based on the Rad-score. Using multivariate logistic regression analysis, the Rad-score and significant predictors were integrated into a single predictive model (radiomics nomogram model 1) to predict the durable clinical benefit (DCB) of ICIs. Radiomics model 2 was developed based on the same Rad-score as radiomics model 1.Using multivariate Cox proportional hazards regression analysis, the Rad-score, and independent risk factors, radiomics nomogram model 2 was constructed to predict the progression-free survival (PFS). RESULTS: The models successfully predicted the patients who would benefit from ICIs. For radiomics model 1, the area under the receiver operating characteristic curve values for the training and validation cohorts were 0.848 and 0.795, respectively, whereas for radiomics nomogram model 1, the values were 0.902 and 0.877, respectively. For the PFS prediction, the Harrell's concordance indexes for the training and validation cohorts were 0.717 and 0.760, respectively, using radiomics model 2, whereas they were 0.749 and 0.791, respectively, using radiomics nomogram model 2. CONCLUSIONS: CT-based radiomic features and clinicopathological factors can be used prior to the initiation of immunotherapy for identifying NSCLC patients who are the most likely to benefit from the therapy. This could guide the individualized treatment strategy for advanced NSCLC.

28Delta-radiomics features combined with haematological index predict pathological complete response after neoadjuvant immunochemotherapy in resectable non-small cell lung cancer.PubMed

D Xiong, J Li, L Li, et al.
Clin Radiol. 2025 Jul;86:106906. doi: 10.1016/j.crad.2025.106906. Epub 2025 Apr 7.
AIM: This study aimed at assessing the value of enhanced computed tomography (CT)-based delta-radiomics features (Delta-RFs) and Delta-RFs combined with haematological dynamic changes in predicting pathological complete response (PCR) after neoadjuvant immunochemotherapy in non-small cell lung cancer (NSCLC). MATERIALS AND METHODS: From January 2021 to August 2023, in total, 165 patients with stage IB-IIIB NSCLC (training, n=115, validation, n=50) who received neoadjuvant immunochemotherapy before surgery, were retrospectively enrolled. Radiomic features were extracted from tumour region of interest on pretreatment and pre-operation enhanced CT images. Delta-RFs are defined as the relative net change in radiomics features between pre-neoadjuvant immunochemotherapy and pre-operation stage. The least absolute shrinkage and selection operator was used to ensure optimal feature selection to calculate the radiomics score (Rad-score) for predicting PCR. Univariate and multivariate logistic regression analyses were performed to screen the factors related to PCR and predictive models were then constructed. RESULTS: Forty percent patients showed PCR (66/165) after neoadjuvant immunochemotherapy. Nine Delta-RFs were selected as the most predictive factors for PCR. Logistic regression analysis showed that the Rad-score (OR = 8.542, 95% CI: 3.367-21.673, P<0.001) and ΔLMR (OR = 2.637, 95% CI: 1.094-6.359, P=0.031) were independent factors associated with PCR. With respect to predicting PCR, the Delta-RF model and the combined model both achieved satisfactory areas under the curve in the training (area under the curve [AUC]: 0.74, 0.788) and the validation was found to be cohort (AUC: 0.718, 0.737). The calibration curve showed that the predicted value of Delta-RF combined with haematological dynamic change model was in good agreement with the observed value. Decision curve analysis represented that the model exhibits high clinical practicability. CONCLUSIONS: The Delta-RF model based on enhanced CT and the combined model can aid in efficient prediction of PCR after neoadjuvant immunochemotherapy in NSCLC, and the combined model can predict PCR performance better than Delta-RF model alone after neoadjuvant immunochemotherapy.

29Prognostic nomogram combining F-FDG PET/CT radiomics and clinical data for stage III NSCLC survival prediction.PubMed

Yalin Zhang, Yongbin Cui, Huiling Liu, et al.
Sci Rep. 2024 Sep 4;14(1):20557. doi: 10.1038/s41598-024-71003-3.
The aim of this study was to establish and validate the precision of a novel radiomics approach that integrates 18Fluorine-fluorodeoxyglucose (18F-FDG) positron emission tomography (PET)-computed tomography (CT) scan data with clinical information to improve the prognostication of survival rates in patients diagnosed with stage III Non-Small Cell Lung Cancer (NSCLC) who are not candidates for surgery. We evaluated pretreatment F-FDG PET-CT scans from 156 individuals diagnosed with stage III inoperable NSCLC at Shandong Cancer Hospital. These individuals were divided into two groups: a training set comprising 110 patients and an internal validation set consisting of 46 patients. By employing random forest classifier and cox proportional hazards model , we identified and utilized relevant features to create predictive models and a nomogram. The effectiveness of these models was assessed through the use of the receiver operating characteristics(ROC) curves, Kaplan-Meier (KM) curves, and the application of the nomogram. Our findings showed that the combined model, which integrates both clinical and radiomic data, outperformed those based solely on clinical or radiomic features in predicting 3-year overall survival(OS). Furthermore, calibration plots revealed a high level of agreement between predicted and actual survival times. The research successfully established a predictive radiomics model that integrates F-FDG PET/CT imaging with clinical indicators to enhance survival predictions for patients with stage III inoperable NSCLC.

30[F]FDG PET-CT radiomics signature to predict pathological complete response to neoadjuvant chemoimmunotherapy in non-small cell lung cancer: a multicenter study.PubMed

Minglei Yang, Xiaoxiao Li, Chuang Cai, et al.
Eur Radiol. 2024 Jul;34(7):4352-4363. doi: 10.1007/s00330-023-10503-8. Epub 2023 Dec 21.
OBJECTIVES: This study aims to develop and validate a radiomics model based on F-fluorodeoxyglucose positron emission tomography-computed tomography ([F]FDG PET-CT) images to predict pathological complete response (pCR) to neoadjuvant chemoimmunotherapy in non-small cell lung cancer (NSCLC). MATERIALS AND METHODS: One hundred eighty-five patients receiving neoadjuvant chemoimmunotherapy for NSCLC at 5 centers from January 2019 to December 2022 were included and divided into a training cohort and a validation cohort. Radiomics models were constructed via the least absolute shrinkage and selection operator (LASSO) method. The performances of models were evaluated by the area under the receiver operating characteristic curve (AUC). In addition, genetic analyses were conducted to reveal the underlying biological basis of the radiomics score. RESULTS: After the LASSO process, 9 PET-CT radiomics features were selected for pCR prediction. In the validation cohort, the ability of PET-CT radiomics model to predict pCR was shown to have an AUC of 0.818 (95% confidence interval [CI], 0.711, 0.925), which was better than the PET radiomics model (0.728 [95% CI, 0.610, 0.846]), CT radiomics model (0.732 [95% CI, 0.607, 0.857]), and maximum standard uptake value (0.603 [95% CI, 0.473, 0.733]) (p < 0.05). Moreover, a high radiomics score was related to the upregulation of pathways suppressing tumor proliferation and the infiltration of antitumor immune cell. CONCLUSION: The proposed PET-CT radiomics model was capable of predicting pCR to neoadjuvant chemoimmunotherapy in NSCLC patients. CLINICAL RELEVANCE STATEMENT: This study indicated that the generated F-fluorodeoxyglucose positron emission tomography-computed tomography radiomics model could predict pathological complete response to neoadjuvant chemoimmunotherapy, implying the potential of our radiomics model to personalize the neoadjuvant chemoimmunotherapy in lung cancer patients. KEY POINTS: • Recognizing patients potentially benefiting neoadjuvant chemoimmunotherapy is critical for individualized therapy of lung cancer. • [F]FDG PET-CT radiomics could predict pathological complete response to neoadjuvant immunotherapy in non-small cell lung cancer. • [F]FDG PET-CT radiomics model could personalize neoadjuvant chemoimmunotherapy in lung cancer patients.

31Multimodal data integration using machine learning improves risk stratification of high-grade serous ovarian cancer.PubMed

Kevin M Boehm, Emily A Aherne, Lora Ellenson, et al.
Nat Cancer. 2022 Jun;3(6):723-733. doi: 10.1038/s43018-022-00388-9. Epub 2022 Jun 28.
Patients with high-grade serous ovarian cancer suffer poor prognosis and variable response to treatment. Known prognostic factors for this disease include homologous recombination deficiency status, age, pathological stage and residual disease status after debulking surgery. Recent work has highlighted important prognostic information captured in computed tomography and histopathological specimens, which can be exploited through machine learning. However, little is known about the capacity of combining features from these disparate sources to improve prediction of treatment response. Here, we assembled a multimodal dataset of 444 patients with primarily late-stage high-grade serous ovarian cancer and discovered quantitative features, such as tumor nuclear size on staining with hematoxylin and eosin and omental texture on contrast-enhanced computed tomography, associated with prognosis. We found that these features contributed complementary prognostic information relative to one another and clinicogenomic features. By fusing histopathological, radiologic and clinicogenomic machine-learning models, we demonstrate a promising path toward improved risk stratification of patients with cancer through multimodal data integration.

32Multimodal integration of radiology, pathology and genomics for prediction of response to PD-(L)1 blockade in patients with non-small cell lung cancer.PubMed

Rami S Vanguri, Jia Luo, Andrew T Aukerman, et al.
Nat Cancer. 2022 Oct;3(10):1151-1164. doi: 10.1038/s43018-022-00416-8. Epub 2022 Aug 29.
Immunotherapy is used to treat almost all patients with advanced non-small cell lung cancer (NSCLC); however, identifying robust predictive biomarkers remains challenging. Here we show the predictive capacity of integrating medical imaging, histopathologic and genomic features to predict immunotherapy response using a cohort of 247 patients with advanced NSCLC with multimodal baseline data obtained during diagnostic clinical workup, including computed tomography scan images, digitized programmed death ligand-1 immunohistochemistry slides and known outcomes to immunotherapy. Using domain expert annotations, we developed a computational workflow to extract patient-level features and used a machine-learning approach to integrate multimodal features into a risk prediction model. Our multimodal model (area under the curve (AUC) = 0.80, 95% confidence interval (CI) 0.74-0.86) outperformed unimodal measures, including tumor mutational burden (AUC = 0.61, 95% CI 0.52-0.70) and programmed death ligand-1 immunohistochemistry score (AUC = 0.73, 95% CI 0.65-0.81). Our study therefore provides a quantitative rationale for using multimodal features to improve prediction of immunotherapy response in patients with NSCLC using expert-guided machine learning.

33Non-invasive multimodal CT deep learning biomarker to predict pathological complete response of non-small cell lung cancer following neoadjuvant immunochemotherapy: a multicenter study.PubMed

Guanchao Ye, Guangyao Wu, Yu Qi, et al.
J Immunother Cancer. 2024 Sep 3;12(9):e009348. doi: 10.1136/jitc-2024-009348.
OBJECTIVES: Although neoadjuvant immunochemotherapy has been widely applied in non-small cell lung cancer (NSCLC), predicting treatment response remains a challenge. We used pretreatment multimodal CT to explore deep learning-based immunochemotherapy response image biomarkers. METHODS: This study retrospectively obtained non-contrast enhanced and contrast enhancedbubu CT scans of patients with NSCLC who underwent surgery after receiving neoadjuvant immunochemotherapy at multiple centers between August 2019 and February 2023. Deep learning features were extracted from both non-contrast enhanced and contrast enhanced CT scans to construct the predictive models (LUNAI-uCT model and LUNAI-eCT model), respectively. After the feature fusion of these two types of features, a fused model (LUNAI-fCT model) was constructed. The performance of the model was evaluated using the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, specificity, positive predictive value, and negative predictive value. SHapley Additive exPlanations analysis was used to quantify the impact of CT imaging features on model prediction. To gain insights into how our model makes predictions, we employed Gradient-weighted Class Activation Mapping to generate saliency heatmaps. RESULTS: The training and validation datasets included 113 patients from Center A at the 8:2 ratio, and the test dataset included 112 patients (Center B n=73, Center C n=20, Center D n=19). In the test dataset, the LUNAI-uCT, LUNAI-eCT, and LUNAI-fCT models achieved AUCs of 0.762 (95% CI 0.654 to 0.791), 0.797 (95% CI 0.724 to 0.844), and 0.866 (95% CI 0.821 to 0.883), respectively. CONCLUSIONS: By extracting deep learning features from contrast enhanced and non-contrast enhanced CT, we constructed the LUNAI-fCT model as an imaging biomarker, which can non-invasively predict pathological complete response in neoadjuvant immunochemotherapy for NSCLC.

34Multimodal integration using a machine learning approach facilitates risk stratification in HR+/HER2- breast cancer.PubMed

Hang Zhang, Fan Yang, Ying Xu, et al.
Cell Rep Med. 2025 Feb 18;6(2):101924. doi: 10.1016/j.xcrm.2024.101924. Epub 2025 Jan 22.
Hormone receptor-positive (HR+)/human epidermal growth factor receptor 2-negative (HER2-) breast cancer is the most common type of breast cancer, with continuous recurrence remaining an important clinical issue. Current relapse predictive models in HR+/HER2- breast cancer patients still have limitations. The integration of multidimensional data represents a promising alternative for predicting relapse. In this study, we leverage our multi-omics cohort comprising 579 HR+/HER2- breast cancer patients (200 patients with complete data across 7 modalities) and develop a machine-learning-based model, namely CIMPTGV, which integrates clinical information, immunohistochemistry, metabolomics, pathomics, transcriptomics, genomics, and copy number variations to predict recurrence risk of HR+/HER2- breast cancer. This model achieves concordance indices (C-indices) of 0.871 and 0.869 in the train and test sets, respectively. The risk population predicted by the CIMPTGV model encompasses those identified by single-modality models. Feature analysis reveals that synergistic and complementary effects exist in different modalities. Simultaneously, we develop a simplified model with a mean area under the curve (AUC) of 0.840, presenting a useful approach for clinical applications.

35A multimodal and fully automated system for prediction of pathological complete response to neoadjuvant chemotherapy in breast cancer.PubMed

Ning Mao, Yi Dai, Heng Zhou, et al.
Sci Adv. 2025 May 2;11(18):eadr1576. doi: 10.1126/sciadv.adr1576. Epub 2025 Apr 30.
Accurately predicting pathological complete response (pCR) before neoadjuvant chemotherapy (NAC) is crucial for patients with breast cancer. In this study, we developed a multimodal integrated fully automated pipeline system (MIFAPS) in forecasting pCR to NAC, using a multicenter and prospective dataset of 1004 patients with locally advanced breast cancer, incorporating pretreatment magnetic resonance imaging, whole slide image, and clinical risk factors. The results demonstrated that MIFAPS offered a favorable predictive performance in both the pooled external test set [area under the curve (AUC) = 0.882] and the prospective test set (AUC = 0.909). In addition, MIFAPS significantly outperformed single-modality models ( < 0.05). Furthermore, the high deep learning scores were associated with immune-related pathways and the promotion of antitumor cells in the microenvironment during biological basis exploration. Overall, our study demonstrates a promising approach for improving the prediction of pCR to NAC in patients with breast cancer through the integration of multimodal data.

36Predictive Value of Metabolic Parameters Derived From F-FDG PET/CT for Microsatellite Instability in Patients With Colorectal Carcinoma.PubMed

Hao Liu, Zheng Ye, Ting Yang, et al.
Front Immunol. 2021 Aug 26;12:724464. doi: 10.3389/fimmu.2021.724464. eCollection 2021.
BACKGROUND: Microsatellite instability (MSI) is one of the important factors that determine the effectiveness of immunotherapy in colorectal cancer (CRC) and serves as a prognostic biomarker for its clinical outcomes. PURPOSE: To investigate whether the metabolic parameters derived fromF-fluorodeoxyglucose (F-FDG) positron emission tomography/computed tomography (PET/CT) can predict MSI status in patients with CRC. MATERIALS AND METHODS: A retrospective analysis was performed on CRC patients who underwent F-FDG PET/CT examination before surgery between January 2015 and April 2021. The metabolic F-FDG PET/CT parameters of the primary CRC lesion were calculated and recorded with different thresholds, including the maximum, peak, and mean standardized uptake value (SUV, SUV, and SUV), as well as the metabolic tumor volume (MTV) and the total lesion glycolysis (TLG). The status of MSI was determined by immunohistochemical assessment. The difference of quantitative parameters between MSI and microsatellite stability (MSS) groups was assessed, and the receiver operating characteristic (ROC) analyses with area under ROC curves (AUC) was used to evaluate the predictive performance of metabolic parameters. RESULTS: A total of 44 patients (24 men and 20 women; mean ± standard deviation age: 71.1 ± 14.2 years) were included. There were 14 patients in the MSI group while there were 30 in the MSS group. MTV, MTV, MTV, and MTV, as well as TLG and TLG showed significant difference between two groups (all -values <0.05), among which MTV demonstrated the highest performance in the prediction of MSI, with an AUC of 0.805 [95% confidence interval (CI): 0.657-0.909], a sensitivity of 92.9% (95% CI: 0.661-0.998), and a specificity of 66.7% (95% CI: 0.472-0.827). Patients' age and MTV were significant predictive indicators of MSI in multivariate logistic regression. CONCLUSION: The metabolic parameters derived fromF-FDG PET/CT were able to preoperatively predict the MSI status in CRC, with MTV demonstrating the highest predictive performance. PET/CT imaging could serve as a noninvasive tool in the guidance of immunotherapy and individualized treatment in CRC patients.

37Circulating tumor HPV DNA in the management of HPV+ oropharyngeal cancer and its correlation with MRI.PubMed

Flaminia Campo, Francesca Paolini, Valentina Manciocco, et al.
Head Neck. 2024 Sep;46(9):2206-2213. doi: 10.1002/hed.27866. Epub 2024 Jul 9.
BACKGROUND: First aim was to compare ddPCR assays of ctHPVDNA with p16 IHC and qualitative HPV PCR. Second aim was to carry out longitudinal blood sampling to test for association of ctHPVDNA with histological confirmed recurrence. Third aim was to perform a multidimensional assessment which included: (1) clinical features; (2) ctHPVDNA; (3) MRI-based tumor size measurements of primary tumor (PT) and cervical lymph node metastases (CLNM). METHODS: Plasma samples were collected before treatment and during follow-up, and ddPCR assay comprising E6 of HPV16 and HPV 33 and HPV 35 was used. RESULTS: Present study was conducted at diagnosis in 117 patients and revealed a ctHPVDNA sensitivity of 100% (95% CI 95.5-100) and a specificity of 94.4 (95% CI 81.3-99.3), positive predictive value (PPV) of 94.4 (95% CI 81.3-99.3), and negative predictive value (NPP) of 100% (95% CI 89.7-100). During follow-up ctHPVDNA had a sensitivity of 100% (95% CI 72.1-100)% and specificity of 98.4% (95% CI 91.7-100)%, PPV% of 90.9% (95% CI 62.3-98.4) and NPV% of 100% (95% CI 94.3-100) for ability to detect recurrence. Correlation between both the CLNM volume and the sum of PT and CLNM volume was observed. CONCLUSIONS: ctHPVDNA was superior to p16 in identification of HPV-OPSCC at diagnosis. Introduction of ctHPVDNA, beyond diagnostic setting, represents a great opportunity to improve follow-up protocol of OPSCC patients.

38Personalized monitoring of circulating tumor DNA with a specific signature of trackable mutations after chimeric antigen receptor T-cell therapy in follicular lymphoma patients.PubMed

Ana Jiménez-Ubieto, Alejandro Martín-Muñoz, María Poza, et al.
Front Immunol. 2023 Jun 5;14:1188818. doi: 10.3389/fimmu.2023.1188818. eCollection 2023.
BACKGROUND: CART therapy has produced a paradigm shift in the treatment of relapsing FL patients. Strategies to optimize disease surveillance after these therapies are increasingly necessary. This study explores the potential value of ctDNA monitoring with an innovative signature of personalized trackable mutations. METHOD: Eleven FL patients treated with anti-CD19 CAR T-cell therapy were included. One did not respond and was excluded. Genomic profiling was performed before starting lymphodepleting chemotherapy to identify somatic mutations suitable for LiqBio-MRD monitoring. The dynamics of the baseline mutations (4.5 per patient) were further analyzed on 59 cfDNA follow-up samples. PET/CT examinations were performed on days +90, +180, +365, and every six months until disease progression or death. RESULTS: After a median follow-up of 36 months, all patients achieved a CR as the best response. Two patients progressed. The most frequently mutated genes were CREBBP, KMT2D and EP300. Simultaneous analysis of ctDNA and PET/CT was available for 18 time-points. When PET/CT was positive, two out of four ctDNA samples were LiqBio-MRD negative. These two negative samples corresponded to women with a unique mesenteric mass in two evaluations and never relapsed. Meanwhile, 14 PET/CT negative images were mutation-free based on our LiqBio-MRD analysis (100%). None of the patients had a negative LiqBio-MRD test by day +7. Interestingly, all durably responding patients had undetectable ctDNA at or around three months after infusion. Two patients presented discordant results by PET/CT and ctDNA levels. No progression was confirmed in these cases. All the progressing patients were LiqBio-MRD positive before progression. CONCLUSION: This is a proof-of-principle for using ctDNA to monitor response to CAR T-cell therapy in FL. Our results confirm that a non-invasive liquid biopsy MRD analysis may correlate with response and could be used to monitor response. Harmonized definitions of ctDNA molecular response and pinpointing the optimal timing for assessing ctDNA responses are necessary for this setting. If using ctDNA analysis, we suggest restricting follow-up PET/CT in CR patients to a clinical suspicion of relapse, to avoid false-positive results.

39Circulating Biomarkers Predictive of Treatment Response in Patients with Hormone-sensitive or Castration-resistant Metastatic Prostate Cancer: A Systematic Review.PubMed

Michael Baboudjian, Arthur Peyrottes, Charles Dariane, et al.
Eur Urol Oncol. 2024 Dec;7(6):1228-1245. doi: 10.1016/j.euo.2024.05.003. Epub 2024 May 31.
BACKGROUND AND OBJECTIVE: Metastatic prostate cancer (mPCa) harbors genomic alterations that may predict targeted therapy efficacy. These alterations can be identified not only in tissue but also directly in biologic fluids (ie, liquid biopsies), mainly blood. Liquid biopsies may represent a safer and less invasive alternative for monitoring patients treated for mPCa. Current research focuses on the description and validation of novel predictive biomarkers to improve precision medicine in mPCa. Our aim was to systematically review the current evidence on liquid biopsy biomarkers for predicting treatment response in mPCa. METHODS: We systematically searched Medline, Web of Science, and evidence-based websites for publications on circulating biomarkers in mPCa between March 2013 and February 2024 for review. Endpoints were: prediction of overall survival, biochemical or radiographic progression-free survival after treatment (chemotherapy, androgen deprivation therapy, androgen receptor pathway inhibitors [ARPIs], immunotherapy, or PARP inhibitors [PARPIs]). For each biomarker, the level of evidence (LOE) for clinical validity was attributed: LOE IA and IB, high level of evidence; LOE IIB and IIC, intermediate level; and LOE IIIC and LOE IV-VD, weak level. KEY FINDINGS AND LIMITATIONS: The predictive value of each biomarker for the response to several therapies was evaluated in both metastatic hormone-sensitive (mHSPC) and castration-resistant prostate cancer (mCRPC). In patients with mCRPC, BRCA1/2 or ATM mutations predicted response to ARPIs (LOE IB) and PARPIs (LOE IIB), while AR-V7 transcripts or AR-V7 protein levels in circulating tumor cells (CTCs) predicted response to ARPIs and taxanes (LOE IB). CTC quantification predicted response to cabazitaxel, abiraterone, and radium-223 (LOE IIB), while TP53 alterations predicted response to Lu prostate-specific membrane antigen radioligand treatment (LOE IIB). AR copy number in circulating tumor DNA before the first treatment line and before subsequent lines predicted response to docetaxel, cabazitaxel, and ARPIs (LOE IIB). In mHSPC, DNA damage in lymphocytes was predictive of the response to radium-223 (LOE IIB). CONCLUSIONS AND CLINICAL IMPLICATIONS: BRCA1/2, ATM, and AR alterations detected in liquid biopsies may help clinicians in management of patients with mPCa. The other circulating biomarkers did not reach the LOE required for routine clinical use and should be validated in prospective independent studies. PATIENT SUMMARY: We reviewed studies assessing the value of biomarkers in blood or urine for management of metastatic prostate cancer. The evidence indicates that some biomarkers could help in selecting patients eligible for specific treatments.

40Molecular Disease Monitoring in Patients With Relapsed/refractory B-Cell Lymphoma Receiving Anti-CD19 CAR T-Cell Therapy.PubMed

Meryl Colton, Enkhtsetseg Purev, Bradley Haverkos, et al.
Clin Lymphoma Myeloma Leuk. 2024 Nov;24(11):778-782. doi: 10.1016/j.clml.2024.06.006. Epub 2024 Jun 26.
BACKGROUND: Chimeric antigen receptor (CAR) T-cell therapy has improved the historically poor outcomes for relapsed and refractory (R/R) large B-cell non-Hodgkin's lymphoma (LBCL). However, nearly 60% of patients will either fail to respond or relapse after CAR T-cell therapy. Currently, PET/CT scans are used to assess response. Cell-free circulating tumor DNA (ctDNA) is released by tumor cells into the peripheral blood and can be measured for minimal residual disease (MRD) assessment. METHODS: In this retrospective, IRB approved pilot study, archived lymphoma tissue and ctDNA from peripheral blood samples on day 0, 14, 28, 56, 90, 180, and 365 after CAR T-cell infusion from 10 patients with R/R NHL were collected for next-generation sequencing (NGS) of clonal variable-diversity-joining (VDJ) rearrangements (Adaptive biotechnologies [Seattle, WA]). Response was assessed by PET/CT on days 90 and 365 and graded according to the Lugano 2014 criteria. The primary endpoint was to determine the feasibility of detecting ctDNA to monitor disease response after anti-CD19 CAR T-cell therapy. The secondary endpoint was to compare the sensitivity/specificity of MRD assessment from ctDNA to PET/CT imaging. RESULTS: Nine out of 10 patients with a trackable sequence [median age 69 (range: 56-76); 55.6% male; median LDH 224], were included in this study. Each received tisagenlecleucel (tisa-cel) CAR T-cell therapy after median 2 prior treatments (range: 2-4). 7/9 patients had R/R diffuse large B-cell lymphoma (DLBCL), and 2/9 had transformed follicular lymphoma. At a median follow up of 12.7 months (range: 1.5-30 months), 4 patients were alive. By day 90, 3 patients (33.3%) achieved a radiographic complete response (CR) whilst 6 patients (66.6%) had progressive disease (PD). Detectable MRD on day 14 or day 28 had 83% sensitivity and 100% specificity for radiographic progression at any time before 1 year. For patients with PD, the median (interquartile range) MRD at day 0, 14, and 28 were 17.31 (1.01, 96.84), 9.12 (0.30, 18.8), and 23.77 (8.01, 137.53) copies per milliliter (mL), respectively. For patients with detectable MRD at day 28, mOS and mPFS were 6.7 and 1.3 months, respectively. CONCLUSION: Monitoring MRD was a sensitive and specific method to detect poor response to tisa-cel. Additional studies evaluating MRD more frequently and with different products are warranted.

41Enhancing NSCLC recurrence prediction with PET/CT habitat imaging, ctDNA, and integrative radiogenomics-blood insights.PubMed

Sheeba J Sujit, Muhammad Aminu, Tatiana V Karpinets, et al.
Nat Commun. 2024 Apr 11;15(1):3152. doi: 10.1038/s41467-024-47512-0.
While we recognize the prognostic importance of clinicopathological measures and circulating tumor DNA (ctDNA), the independent contribution of quantitative image markers to prognosis in non-small cell lung cancer (NSCLC) remains underexplored. In our multi-institutional study of 394 NSCLC patients, we utilize pre-treatment computed tomography (CT) and F-fluorodeoxyglucose positron emission tomography (FDG-PET) to establish a habitat imaging framework for assessing regional heterogeneity within individual tumors. This framework identifies three PET/CT subtypes, which maintain prognostic value after adjusting for clinicopathologic risk factors including tumor volume. Additionally, these subtypes complement ctDNA in predicting disease recurrence. Radiogenomics analysis unveil the molecular underpinnings of these imaging subtypes, highlighting downregulation in interferon alpha and gamma pathways in the high-risk subtype. In summary, our study demonstrates that these habitat imaging subtypes effectively stratify NSCLC patients based on their risk levels for disease recurrence after initial curative surgery or radiotherapy, providing valuable insights for personalized treatment approaches.

42Integration of multi-omics data for survival prediction of lung adenocarcinoma.PubMed

Dingjie Guo, Yixian Wang, Jing Chen, et al.
Comput Methods Programs Biomed. 2024 Jun;250:108192. doi: 10.1016/j.cmpb.2024.108192. Epub 2024 Apr 22.
BACKGROUND AND OBJECTIVE: The morbidity of lung adenocarcinoma (LUAD) has been increasing year by year and the prognosis is poor. This has prompted researchers to study the survival of LUAD patients to ensure that patients can be cured in time or survive after appropriate treatment. There is still no fully valid model that can be applied to clinical practice. METHODS: We introduced struc2vec-based multi-omics data integration (SBMOI), which could integrate gene expression, somatic mutations and clinical data to construct mutation gene vectors representing LUAD patient features. Based on the patient features, the random survival forest (RSF) model was used to predict the long- and short-term survival of LUAD patients. To further demonstrate the superiority of SBMOI, we simultaneously replaced scale-free gene co-expression network (FCN) with a protein-protein interaction (PPI) network and a significant co-expression network (SCN) to compare accuracy in predicting LUAD patient survival under the same conditions. RESULTS: Our results suggested that compared with SCN and PPI network, the FCN based SBMOI combined with RSF model had better performance in long- and short-term survival prediction tasks for LUAD patients. The AUC of 1-year, 5-year, and 10-year survival in the validation dataset were 0.791, 0.825, and 0.917, respectively. CONCLUSIONS: This study provided a powerful network-based method to multi-omics data integration. SBMOI combined with RSF successfully predicted long- and short-term survival of LUAD patients, especially with high accuracy on long-term survival. Besides, SBMOI algorithm has the potential to combine with other machine learning models to complete clustering or stratificational tasks, and being applied to other diseases.

43Integrated machine learning identifies epithelial cell marker genes for improving outcomes and immunotherapy in prostate cancer.PubMed

Weian Zhu, Hengda Zeng, Jiongduan Huang, et al.
J Transl Med. 2023 Nov 4;21(1):782. doi: 10.1186/s12967-023-04633-2.
BACKGROUND: Prostate cancer (PCa), a globally prevalent malignancy, displays intricate heterogeneity within its epithelial cells, closely linked with disease progression and immune modulation. However, the clinical significance of genes and biomarkers associated with these cells remains inadequately explored. To address this gap, this study aimed to comprehensively investigate the roles and clinical value of epithelial cell-related genes in PCa. METHODS: Leveraging single-cell sequencing data from GSE176031, we conducted an extensive analysis to identify epithelial cell marker genes (ECMGs). Employing consensus clustering analysis, we evaluated the correlations between ECMGs, prognosis, and immune responses in PCa. Subsequently, we developed and validated an optimal prognostic signature, termed the epithelial cell marker gene prognostic signature (ECMGPS), through synergistic analysis from 101 models employing 10 machine learning algorithms across five independent cohorts. Additionally, we collected clinical features and previously published signatures from the literature for comparative analysis. Furthermore, we explored the clinical utility of ECMGPS in immunotherapy and drug selection using multi-omics analysis and the IMvigor cohort. Finally, we investigated the biological functions of the hub gene, transmembrane p24 trafficking protein 3 (TMED3), in PCa using public databases and experiments. RESULTS: We identified a comprehensive set of 543 ECMGs and established a strong correlation between ECMGs and both the prognostic evaluation and immune classification in PCa. Notably, ECMGPS exhibited robust predictive capability, surpassing traditional clinical features and 80 published signatures in terms of both independence and accuracy across five cohorts. Significantly, ECMGPS demonstrated significant promise in identifying potential PCa patients who might benefit from immunotherapy and personalized medicine, thereby moving us nearer to tailored therapeutic approaches for individuals. Moreover, the role of TMED3 in promoting malignant proliferation of PCa cells was validated. CONCLUSIONS: Our findings highlight ECMGPS as a powerful tool for improving PCa patient outcomes and supply a robust conceptual framework for in-depth examination of PCa complexities. Simultaneously, our study has the potential to develop a novel alternative for PCa diagnosis and prognostication.

44Novel tools for early diagnosis and precision treatment based on artificial intelligence.PubMed

Jun Shao, Jiaming Feng, Jingwei Li, et al.
Chin Med J Pulm Crit Care Med. 2023 Sep 9;1(3):148-160. doi: 10.1016/j.pccm.2023.05.001. eCollection 2023 Sep.
Lung cancer has the highest mortality rate among all cancers in the world. Hence, early diagnosis and personalized treatment plans are crucial to improving its 5-year survival rate. Chest computed tomography (CT) serves as an essential tool for lung cancer screening, and pathology images are the gold standard for lung cancer diagnosis. However, medical image evaluation relies on manual labor and suffers from missed diagnosis or misdiagnosis, and physician heterogeneity. The rapid development of artificial intelligence (AI) has brought a whole novel opportunity for medical task processing, demonstrating the potential for clinical application in lung cancer diagnosis and treatment. AI technologies, including machine learning and deep learning, have been deployed extensively for lung nodule detection, benign and malignant classification, and subtype identification based on CT images. Furthermore, AI plays a role in the non-invasive prediction of genetic mutations and molecular status to provide the optimal treatment regimen, and applies to the assessment of therapeutic efficacy and prognosis of lung cancer patients, enabling precision medicine to become a reality. Meanwhile, histology-based AI models assist pathologists in typing, molecular characterization, and prognosis prediction to enhance the efficiency of diagnosis and treatment. However, the leap to extensive clinical application still faces various challenges, such as data sharing, standardized label acquisition, clinical application regulation, and multimodal integration. Nevertheless, AI holds promising potential in the field of lung cancer to improve cancer care.

45Integration of deep learning-based image analysis and genomic data in cancer pathology: A systematic review.PubMed

Lucas Schneider, Sara Laiouar-Pedari, Sara Kuntz, et al.
Eur J Cancer. 2022 Jan;160:80-91. doi: 10.1016/j.ejca.2021.10.007. Epub 2021 Nov 19.
BACKGROUND: Over the past decade, the development of molecular high-throughput methods (omics) increased rapidly and provided new insights for cancer research. In parallel, deep learning approaches revealed the enormous potential for medical image analysis, especially in digital pathology. Combining image and omics data with deep learning tools may enable the discovery of new cancer biomarkers and a more precise prediction of patient prognosis. This systematic review addresses different multimodal fusion methods of convolutional neural network-based image analyses with omics data, focussing on the impact of data combination on the classification performance. METHODS: PubMed was screened for peer-reviewed articles published in English between January 2015 and June 2021 by two independent researchers. Search terms related to deep learning, digital pathology, omics, and multimodal fusion were combined. RESULTS: We identified a total of 11 studies meeting the inclusion criteria, namely studies that used convolutional neural networks for haematoxylin and eosin image analysis of patients with cancer in combination with integrated omics data. Publications were categorised according to their endpoints: 7 studies focused on survival analysis and 4 studies on prediction of cancer subtypes, malignancy or microsatellite instability with spatial analysis. CONCLUSIONS: Image-based classifiers already show high performances in prognostic and predictive cancer diagnostics. The integration of omics data led to improved performance in all studies described here. However, these are very early studies that still require external validation to demonstrate their generalisability and robustness. Further and more comprehensive studies with larger sample sizes are needed to evaluate performance and determine clinical benefits.

46An Appraisal of Lung Nodules Automatic Classification Algorithms for CT Images.PubMed

Xinqi Wang, Keming Mao, Lizhe Wang, et al.
Sensors (Basel). 2019 Jan 7;19(1):194. doi: 10.3390/s19010194.
Lung cancer is one of the most deadly diseases around the world representing about 26% of all cancers in 2017. The five-year cure rate is only 18% despite great progress in recent diagnosis and treatment. Before diagnosis, lung nodule classification is a key step, especially since automatic classification can help clinicians by providing a valuable opinion. Modern computer vision and machine learning technologies allow very fast and reliable CT image classification. This research area has become very hot for its high efficiency and labor saving. The paper aims to draw a systematic review of the state of the art of automatic classification of lung nodules. This research paper covers published works selected from the Web of Science, IEEEXplore, and DBLP databases up to June 2018. Each paper is critically reviewed based on objective, methodology, research dataset, and performance evaluation. Mainstream algorithms are conveyed and generic structures are summarized. Our work reveals that lung nodule classification based on deep learning becomes dominant for its excellent performance. It is concluded that the consistency of the research objective and integration of data deserves more attention. Moreover, collaborative works among developers, clinicians, and other parties should be strengthened.

47Role of artificial intelligence in digital pathology for gynecological cancers.PubMed

Ya-Li Wang, Song Gao, Qian Xiao, et al.
Comput Struct Biotechnol J. 2024 Mar 11;24:205-212. doi: 10.1016/j.csbj.2024.03.007. eCollection 2024 Dec.
The diagnosis of cancer is typically based on histopathological sections or biopsies on glass slides. Artificial intelligence (AI) approaches have greatly enhanced our ability to extract quantitative information from digital histopathology images as a rapid growth in oncology data. Gynecological cancers are major diseases affecting women's health worldwide. They are characterized by high mortality and poor prognosis, underscoring the critical importance of early detection, treatment, and identification of prognostic factors. This review highlights the various clinical applications of AI in gynecological cancers using digitized histopathology slides. Particularly, deep learning models have shown promise in accurately diagnosing, classifying histopathological subtypes, and predicting treatment response and prognosis. Furthermore, the integration with transcriptomics, proteomics, and other multi-omics techniques can provide valuable insights into the molecular features of diseases. Despite the considerable potential of AI, substantial challenges remain. Further improvements in data acquisition and model optimization are required, and the exploration of broader clinical applications, such as the biomarker discovery, need to be explored.

48Deep Learning in Omics Data Analysis and Precision MedicinePubMed

Jordi Martorell-Marugán, Siham Tabik, Yassir Benhammou, et al.
The rise of omics techniques has resulted in an explosion of molecular data in modern biomedical research. Together with information from medical images and clinical data, the field of omics has driven the implementation of personalized medicine. Biomedical and omics datasets are complex and heterogeneous, and extracting meaningful knowledge from this vast amount of information is by far the most important challenge for bioinformatics and machine learning researchers. In this context, there is an increasing interest in the potential of deep learning (DL) methods to create predictive models and to identify complex patterns from these large datasets. This chapter provides an overview of the main applications of DL methods in biomedical research, with focus on omics data analysis and precision medicine applications. DL algorithms and the most popular architectures are introduced first. This is followed by a review of some of the main applications and problems approached by DL in omics data and medical image analysis. Finally, implementations for improving the diagnosis, treatment, and classification of complex diseases are discussed.

49Long-term cancer survival prediction using multimodal deep learning.PubMed

Luís A Vale-Silva, Karl Rohr
Sci Rep. 2021 Jun 29;11(1):13505. doi: 10.1038/s41598-021-92799-4.
The age of precision medicine demands powerful computational techniques to handle high-dimensional patient data. We present MultiSurv, a multimodal deep learning method for long-term pan-cancer survival prediction. MultiSurv uses dedicated submodels to establish feature representations of clinical, imaging, and different high-dimensional omics data modalities. A data fusion layer aggregates the multimodal representations, and a prediction submodel generates conditional survival probabilities for follow-up time intervals spanning several decades. MultiSurv is the first non-linear and non-proportional survival prediction method that leverages multimodal data. In addition, MultiSurv can handle missing data, including single values and complete data modalities. MultiSurv was applied to data from 33 different cancer types and yields accurate pan-cancer patient survival curves. A quantitative comparison with previous methods showed that Multisurv achieves the best results according to different time-dependent metrics. We also generated visualizations of the learned multimodal representation of MultiSurv, which revealed insights on cancer characteristics and heterogeneity.

50BioFusionNet: Deep Learning-Based Survival Risk Stratification in ER+ Breast Cancer Through Multifeature and Multimodal Data Fusion.PubMed

Raktim Kumar Mondol, Ewan K A Millar, Arcot Sowmya, et al.
IEEE J Biomed Health Inform. 2024 Sep;28(9):5290-5302. doi: 10.1109/JBHI.2024.3418341. Epub 2024 Sep 5.
Breast cancer is a significant health concern affecting millions of women worldwide. Accurate survival risk stratification plays a crucial role in guiding personalised treatment decisions and improving patient outcomes. Here we present BioFusionNet, a deep learning framework that fuses image-derived features with genetic and clinical data to obtain a holistic profile and achieve survival risk stratification of ER+ breast cancer patients. We employ multiple self-supervised feature extractors (DINO and MoCoV3) pretrained on histopathological patches to capture detailed image features. These features are then fused by a variational autoencoder and fed to a self-attention network generating patient-level features. A co-dual-cross-attention mechanism combines the histopathological features with genetic data, enabling the model to capture the interplay between them. Additionally, clinical data is incorporated using a feed-forward network, further enhancing predictive performance and achieving comprehensive multimodal feature integration. Furthermore, we introduce a weighted Cox loss function, specifically designed to handle imbalanced survival data, which is a common challenge. Our model achieves a mean concordance index of 0.77 and a time-dependent area under the curve of 0.84, outperforming state-of-the-art methods. It predicts risk (high versus low) with prognostic significance for overall survival in univariate analysis (HR=2.99, 95% CI: 1.88-4.78, p 0.005), and maintains independent significance in multivariate analysis incorporating standard clinicopathological variables (HR=2.91, 95% CI: 1.80-4.68, p 0.005).

51A lung nodule segmentation model based on the transformer with multiple thresholds and coordinate attention.PubMed

Tianjiao Hu, Yihua Lan, Yingqi Zhang, et al.
Sci Rep. 2024 Dec 30;14(1):31743. doi: 10.1038/s41598-024-82877-8.
Accurate lung nodule segmentation is fundamental for the early detection of lung cancer. With the rapid development of deep learning, lung nodule segmentation models based on the encoder-decoder structure have become the mainstream research approach. However, during the encoding process, most models have limitations in extracting edge and semantic information and in capturing long-range dependencies. To address these problems, we propose a new lung nodule segmentation model, abbreviated as MCAT-Net. In this model, we construct a multi-threshold feature separation module to capture edge and texture features from different levels and specified intensities of the input image. Secondly, we introduce the coordinate attention mechanism, which allows the model to better recognize and utilize spatial information when handling long-range dependencies, enabling the deep network to maintain its sensitivity to nodule positions. Thirdly, we use the transformer to fully capture the long-range dependencies, further enhancing the global information integration of the network. The proposed method was verified on the LIDC-IDRI and LNDb datasets. The Dice similarity coefficient (DSC) values achieved were 88.29% and 78.51%, and the sensitivities were 86.33% and 75.05%, respectively. The experimental results demonstrated its high practical value for the early diagnosis of lung cancer.

52Novel pre-spatial data fusion deep learning approach for multimodal volumetric outcome prediction models in radiotherapy.PubMed

John C Asbach, Anurag K Singh, Austin J Iovoli, et al.
Med Phys. 2025 Apr;52(4):2675-2687. doi: 10.1002/mp.17672. Epub 2025 Feb 10.
BACKGROUND: Given the recent increased emphasis on multimodal neural networks to solve complex modeling tasks, the problem of outcome prediction for a course of treatment can be framed as fundamentally multimodal in nature. A patient's response to treatment will vary based on their specific anatomy and the proposed treatment plan-these factors are spatial and closely related. However, additional factors may also have importance, such as non-spatial descriptive clinical characteristics, which can be structured as tabular data. It is critical to provide models with as comprehensive of a patient representation as possible, but inputs with differing data structures are incompatible in raw form; traditional models that consider these inputs require feature engineering prior to modeling. In neural networks, feature engineering can be organically integrated into the model itself, under one governing optimization, rather than performed prescriptively beforehand. However, the native incompatibility of different data structures must be addressed. Methods to reconcile structural incompatibilities in multimodal model inputs are called data fusion. We present a novel joint early pre-spatial (JEPS) fusion technique and demonstrate that differences in fusion approach can produce significant model performance differences even when the data is identical. PURPOSE: To present a novel pre-spatial fusion technique for volumetric neural networks and demonstrate its impact on model performance for pretreatment prediction of overall survival (OS). METHODS: From a retrospective cohort of 531 head and neck patients treated at our clinic, we prepared an OS dataset of 222 data-complete cases at a 2-year post-treatment time threshold. Each patient's data included CT imaging, dose array, approved structure set, and a tabular summary of the patient's demographics and survey data. To establish single-modality baselines, we fit both a Cox Proportional Hazards model (CPH) and a dense neural network on only the tabular data, then we trained a 3D convolutional neural network (CNN) on only the volume data. Then, we trained five competing architectures for fusion of both modalities: two early fusion models, a late fusion model, a traditional joint fusion model, and the novel JEPS, where clinical data is merged into training upstream of most convolution operations. We used standardized 10-fold cross validation to directly compare the performance of all models on identical train/test splits of patients, using area under the receiver-operator curve (AUC) as the primary performance metric. We used a two-tailed Student t-test to assess the statistical significance (p-value threshold 0.05) of any observed performance differences. RESULTS: The JEPS design scored the highest, achieving a mean AUC of 0.779 ± 0.080. The late fusion model and clinical-only CPH model scored second and third highest with 0.746 ± 0.066 and 0.720 ± 0.091 mean AUC, respectively. The performance differences between these three models were not statistically significant. All other comparison models scored significantly worse than the top performing JEPS model. CONCLUSION: For our OS evaluation, our JEPS fusion architecture achieves better integration of inputs and significantly improves predictive performance over most common multimodal approaches. The JEPS fusion technique is easily applied to any volumetric CNN.

53Comparison of data fusion strategies for automated prostate lesion detection using mpMRI correlated with whole mount histology.PubMed

Deepa Darshini Gunashekar, Lars Bielak, Benedict Oerther, et al.
Radiat Oncol. 2024 Jul 29;19(1):96. doi: 10.1186/s13014-024-02471-0.
BACKGROUND: In this work, we compare input level, feature level and decision level data fusion techniques for automatic detection of clinically significant prostate lesions (csPCa). METHODS: Multiple deep learning CNN architectures were developed using the Unet as the baseline. The CNNs use both multiparametric MRI images (T2W, ADC, and High b-value) and quantitative clinical data (prostate specific antigen (PSA), PSA density (PSAD), prostate gland volume & gross tumor volume (GTV)), and only mp-MRI images (n = 118), as input. In addition, co-registered ground truth data from whole mount histopathology images (n = 22) were used as a test set for evaluation. RESULTS: The CNNs achieved for early/intermediate / late level fusion a precision of 0.41/0.51/0.61, recall value of 0.18/0.22/0.25, an average precision of 0.13 / 0.19 / 0.27, and F scores of 0.55/0.67/ 0.76. Dice Sorensen Coefficient (DSC) was used to evaluate the influence of combining mpMRI with parametric clinical data for the detection of csPCa. We compared the DSC between the predictions of CNN's trained with mpMRI and parametric clinical and the CNN's trained with only mpMRI images as input with the ground truth. We obtained a DSC of data 0.30/0.34/0.36 and 0.26/0.33/0.34 respectively. Additionally, we evaluated the influence of each mpMRI input channel for the task of csPCa detection and obtained a DSC of 0.14 / 0.25 / 0.28. CONCLUSION: The results show that the decision level fusion network performs better for the task of prostate lesion detection. Combining mpMRI data with quantitative clinical data does not show significant differences between these networks (p = 0.26/0.62/0.85). The results show that CNNs trained with all mpMRI data outperform CNNs with less input channels which is consistent with current clinical protocols where the same input is used for PI-RADS lesion scoring. TRIAL REGISTRATION: The trial was registered retrospectively at the German Register for Clinical Studies (DRKS) under proposal number Nr. 476/14 & 476/19.

54Deep learning in cancer genomics and histopathology.PubMed

Michaela Unger, Jakob Nikolas Kather
Genome Med. 2024 Mar 27;16(1):44. doi: 10.1186/s13073-024-01315-6.
Histopathology and genomic profiling are cornerstones of precision oncology and are routinely obtained for patients with cancer. Traditionally, histopathology slides are manually reviewed by highly trained pathologists. Genomic data, on the other hand, is evaluated by engineered computational pipelines. In both applications, the advent of modern artificial intelligence methods, specifically machine learning (ML) and deep learning (DL), have opened up a fundamentally new way of extracting actionable insights from raw data, which could augment and potentially replace some aspects of traditional evaluation workflows. In this review, we summarize current and emerging applications of DL in histopathology and genomics, including basic diagnostic as well as advanced prognostic tasks. Based on a growing body of evidence, we suggest that DL could be the groundwork for a new kind of workflow in oncology and cancer research. However, we also point out that DL models can have biases and other flaws that users in healthcare and research need to know about, and we propose ways to address them.

55Weakly Supervised Deep Learning in Radiology.PubMed

Leo Misera, Gustav Müller-Franzes, Daniel Truhn, et al.
Radiology. 2024 Jul;312(1):e232085. doi: 10.1148/radiol.232085.
Deep learning (DL) is currently the standard artificial intelligence tool for computer-based image analysis in radiology. Traditionally, DL models have been trained with strongly supervised learning methods. These methods depend on reference standard labels, typically applied manually by experts. In contrast, weakly supervised learning is more scalable. Weak supervision comprises situations in which only a portion of the data are labeled (incomplete supervision), labels refer to a whole region or case as opposed to a precisely delineated image region (inexact supervision), or labels contain errors (inaccurate supervision). In many applications, weak labels are sufficient to train useful models. Thus, weakly supervised learning can unlock a large amount of otherwise unusable data for training DL models. One example of this is using large language models to automatically extract weak labels from free-text radiology reports. Here, we outline the key concepts in weakly supervised learning and provide an overview of applications in radiologic image analysis. With more fundamental and clinical translational work, weakly supervised learning could facilitate the uptake of DL in radiology and research workflows by enabling large-scale image analysis and advancing the development of new DL-based biomarkers.

56Structured Transformation of Unstructured Prostate MRI Reports Using Large Language Models.PubMed

Luca Di Palma, Fatemeh Darvizeh, Marco Alì, et al.
Tomography. 2025 Jun 17;11(6):69. doi: 10.3390/tomography11060069.
OBJECTIVES: to assess the ability of high-performing open-weight large language models (LLMs) in extracting key radiological features from prostate MRI reports. METHODS: Five LLMs (Llama3.3, DeepSeek-R1-Llama3.3, Phi4, Gemma-2, and Qwen2.5-14B) were used to analyze free-text MRI reports retrieved from clinical practice. Each LLM processed reports three times using specialized prompts to extract (1) dimensions, (2) volume and PSA density, and (3) lesion characteristics. An experienced radiologist manually annotated the dataset, defining entities (Exam) and sub-entities (Lesion, Dimension). Feature- and physician-level performance were then assessed. RESULTS: 250 MRI exams reported by 7 radiologists were analyzed by the LLMs. Feature-level performances showed that DeepSeek-R1-Llama3.3 exhibited the highest average score (98.6% ± 2.1%), followed by Phi4 (98.1% ± 2.2%), Llama3.3 (98.0% ± 3.0%), Qwen2.5 (97.5% ± 3.9%), and Gemma2 (96.0% ± 3.4%). All models excelled in extracting PSA density (100%) and volume (≥98.4%), while lesions' extraction showed greater variability (88.4-94.0%). LLMs' performance varied among radiologists: Physician B's reports yielded the highest mean score (99.9% ± 0.2%), while Physician C's resulted in the lowest (94.4% ± 2.3%). CONCLUSIONS: LLMs showed promising results in automated feature-extraction from radiology reports, with DeepSeek-R1-Llama3.3 achieving the highest overall score. These models can improve clinical workflows by structuring unstructured medical text. However, a preliminary analysis of reporting styles is necessary to identify potential challenges and optimize prompt design to better align with individual physician reporting styles. This approach can further enhance the robustness and adaptability of LLM-driven clinical data extraction.

57Extraction of clinical data on major pulmonary diseases from unstructured radiologic reports using a large language model.PubMed

Hyung Jun Park, Jin-Young Huh, Ganghee Chae, et al.
PLoS One. 2024 Nov 25;19(11):e0314136. doi: 10.1371/journal.pone.0314136. eCollection 2024.
Despite significant strides in big data technology, extracting information from unstructured clinical data remains a formidable challenge. This study investigated the utility of large language models (LLMs) for extracting clinical data from unstructured radiological reports without additional training. In this retrospective study, 1800 radiologic reports, 600 from each of the three university hospitals, were collected, with seven pulmonary outcomes defined. Three pulmonology-trained specialists discerned the presence or absence of diseases. Data extraction from the reports was executed using Google Gemini Pro 1.0, OpenAI's GPT-3.5, and GPT-4. The gold standard was predicated on agreement between at least two pulmonologists. This study evaluated the performance of the three LLMs in diagnosing seven pulmonary diseases (active tuberculosis, emphysema, interstitial lung disease, lung cancer, pleural effusion, pneumonia, and pulmonary edema) utilizing chest radiography and computed tomography scans. All models exhibited high accuracy (0.85-1.00) for most conditions. GPT-4 consistently outperformed its counterparts, demonstrating a sensitivity of 0.71-1.00; specificity of 0.89-1.00; and accuracy of 0.89 and 0.99 across both modalities, thus underscoring its superior capability in interpreting radiological reports. Notably, the accuracy of pleural effusion and emphysema on chest radiographs and pulmonary edema on chest computed tomography scans reached 0.99. The proficiency of LLMs, particularly GPT-4, in accurately classifying unstructured radiological data hints at their potential as alternatives to the traditional manual chart reviews conducted by clinicians.

58Large Language Models for Diagnosing Focal Liver Lesions From CT/MRI Reports: A Comparative Study With Radiologists.PubMed

Liuji Sheng, Yidi Chen, Hong Wei, et al.
Liver Int. 2025 Jun;45(6):e70115. doi: 10.1111/liv.70115.
BACKGROUND & AIMS: Whether large language models (LLMs) could be integrated into the diagnostic workflow of focal liver lesions (FLLs) remains unclear. We aimed to investigate two generic LLMs (ChatGPT-4o and Gemini) regarding their diagnostic accuracies referring to the CT/MRI reports, compared to and combined with radiologists of different experience levels. METHODS: From April 2022 to April 2024, this single-center retrospective study included consecutive adult patients who underwent contrast-enhanced CT/MRI for single FLL and subsequent histopathologic examination. The LLMs were prompted by clinical information and the "findings" section of radiology reports three times to provide differential diagnoses in the descending order of likelihood, with the first considered the final diagnosis. In the research setting, six radiologists (three junior and three middle-level) independently reviewed the CT/MRI images and clinical information in two rounds (first alone, then with LLM assistance). In the clinical setting, diagnoses were retrieved from the "impressions" section of radiology reports. Diagnostic accuracy was investigated against histopathology. RESULTS: 228 patients (median age, 59 years; 155 males) with 228 FLLs (median size, 3.6 cm) were included. Regarding the final diagnosis, the accuracy of two-step ChatGPT-4o (78.9%) was higher than single-step ChatGPT-4o (68.0%, p < 0.001) and single-step Gemini (73.2%, p = 0.004), similar to real-world radiology reports (80.0%, p = 0.34) and junior radiologists (78.9%-82.0%; p-values, 0.21 to > 0.99), but lower than middle-level radiologists (84.6%-85.5%; p-values, 0.001 to 0.02). No incremental diagnostic value of ChatGPT-4o was observed for any radiologist (p-values, 0.63 to > 0.99). CONCLUSION: Two-step ChatGPT-4o showed matching accuracies to real-world radiology reports and junior radiologists for diagnosing FLLs but was less accurate than middle-level radiologists and demonstrated little incremental diagnostic value.

59Evaluating the performance of large language models: ChatGPT and Google Bard in generating differential diagnoses in clinicopathological conferences of neurodegenerative disorders.PubMed

Shunsuke Koga, Nicholas B Martin, Dennis W Dickson
Brain Pathol. 2024 May;34(3):e13207. doi: 10.1111/bpa.13207. Epub 2023 Aug 8.
This study explores the utility of the large language models (LLMs), specifically ChatGPT and Google Bard, in predicting neuropathologic diagnoses from clinical summaries. A total of 25 cases of neurodegenerative disorders presented at Mayo Clinic brain bank Clinico-Pathological Conferences were analyzed. The LLMs provided multiple pathologic diagnoses and their rationales, which were compared with the final clinical diagnoses made by physicians. ChatGPT-3.5, ChatGPT-4, and Google Bard correctly made primary diagnoses in 32%, 52%, and 40% of cases, respectively, while correct diagnoses were included in 76%, 84%, and 76% of cases, respectively. These findings highlight the potential of artificial intelligence tools like ChatGPT in neuropathology, suggesting they may facilitate more comprehensive discussions in clinicopathological conferences.

60Enhancing Patient-Trial Matching With Large Language Models: A Scoping Review of Emerging Applications and Approaches.PubMed

Hongyu Chen, Xiaohan Li, Xing He, et al.
JCO Clin Cancer Inform. 2025 Jun;9:e2500071. doi: 10.1200/CCI-25-00071. Epub 2025 Jun 9.
PURPOSE: Patient recruitment remains a major bottleneck in clinical trial execution, with inefficient patient-trial matching often causing delays and failures. Recent advancements in large language models (LLMs) offer a promising avenue for automating and improving this process. This scoping review aims to provide a comprehensive synthesis of the emerging applications of LLMs in patient-trial matching. METHODS: A comprehensive search was conducted in PubMed, Web of Science, and OpenAlex for literature published between December 1, 2022, and December 31, 2024. Studies were included if they explicitly integrated LLMs into patient-trial matching systems. Data extraction focused on system architectures, patient data processing, eligibility criteria processing, matching techniques, evaluation metrics, and performance. RESULTS: Of the 2,357 studies initially identified, 24 met the inclusion criteria. The majority (21/24) were published in 2024, highlighting the rapid adoption of LLMs in this domain. Most systems used patient-centric matching (17/24), with OpenAI's generative pretrained transformer models being the most commonly used LLM. Core components of these systems included eligibility criteria processing, patient data processing, and matching, with some incorporating retrieval algorithms to enhance computational efficiency. LLM-integrated approaches demonstrated improved accuracy and scalability in patient-trial matching, although challenges such as performance variability, interpretability, and reliance on synthetic data sets remain significant. CONCLUSION: LLM-based patient-trial matching systems present a transformative opportunity to enhance the efficiency and accuracy of clinical trial recruitment. Despite current limitations related to model generalizability, explainability, and data constraints, future advancements in hybrid modeling strategies, domain-specific fine-tuning, and real-world data set integration could further optimize LLM-based trial matching. Addressing these challenges will be crucial to realizing the full potential of LLMs in streamlining patient recruitment and accelerating clinical trial execution.

61Applying Large Language Models for Surgical Case Length Prediction.PubMed

Adhitya Ramamurthi, Bhabishya Neupane, Priya Deshpande, et al.
JAMA Surg. 2025 Jul 9. doi: 10.1001/jamasurg.2025.2154.
IMPORTANCE: Accurate prediction of surgical case duration is critical for operating room (OR) management, as inefficient scheduling can lead to reduced patient and surgeon satisfaction while incurring considerable financial costs. OBJECTIVE: To evaluate the feasibility and accuracy of large language models (LLMs) in predicting surgical case length using unstructured clinical data compared to existing estimation methods. DESIGN, SETTING, AND PARTICIPANTS: This was a retrospective study analyzing elective surgical cases performed between January 2017 and December 2023 at a single academic medical center and affiliated community hospital ORs. Analysis included 125 493 eligible surgical cases, with 1950 used for LLM fine-tuning and 2500 for evaluation. An additional 500 cases from a community site were used for external validation. Cases were randomly sampled using strata to ensure representation across surgical specialties. EXPOSURES: Eleven LLMs, including base models (GPT-4, GPT-3.5, Mistral, Llama-3, Phi-3) and 2 fine-tuned variants (GPT-4 fine-tuned, GPT-3.5 fine-tuned), were used to predict surgical case length based on clinical notes. MAIN OUTCOMES AND MEASURES: The primary outcome was average error between predicted and actual surgical case length (wheels-in to wheels-out time). The secondary outcome was prediction accuracy, defined as predicted length within 20% of actual duration. RESULTS: Fine-tuned GPT-4 achieved the best performance with a mean absolute error (MAE) of 47.64 minutes (95% CI, 45.71-49.56) and R2 of 0.61, matching the performance of current OR scheduling (MAE, 49.34 minutes; 95% CI, 47.60-51.09; R2, 0.63; P = .10). Both GPT-4 fine-tuned and GPT-3.5 fine-tuned significantly outperformed current scheduling methods in accuracy (46.12% and 46.08% vs 40.92%, respectively; P < .001). GPT-4 fine-tuned outperformed all other models during external validation with similar performance metrics (MAE, 48.66 minutes; 95% CI, 45.31-52.00; accuracy, 46.0%). Base models demonstrated variable performance, with GPT-4 showing the highest performance among non-fine-tuned models (MAE, 59.20 minutes; 95% CI, 56.88 - 61.52). CONCLUSION AND RELEVANCE: The findings in this study suggest that fine-tuned LLMs can predict surgical case length with accuracy comparable to or exceeding current institutional scheduling methods. This indicates potential for LLMs to enhance operating room efficiency through improved case length prediction using existing clinical documentation.

62Model development for bespoke large language models for digital triage assistance in mental health care.PubMed

Niall Taylor, Andrey Kormilitzin, Isabelle Lorge, et al.
Artif Intell Med. 2024 Nov;157:102988. doi: 10.1016/j.artmed.2024.102988. Epub 2024 Sep 29.
Contemporary large language models (LLMs) may have utility for processing unstructured, narrative free-text clinical data contained in electronic health records (EHRs) - a particularly important use-case for mental health where a majority of routinely-collected patient data lacks structured, machine-readable content. A significant problem for the United Kingdom's National Health Service (NHS) are the long waiting lists for specialist mental healthcare. According to NHS data (NHS Digital, 2024), in each month of 2023, there were between 370,000 and 470,000 individual new referrals into secondary mental healthcare services. Referrals must be triaged by clinicians, using clinical information contained in the patient's EHR to arrive at a decision about the most appropriate mental healthcare team to assess and potentially treat these patients. The ability to efficiently recommend a relevant team by ingesting potentially voluminous clinical notes could help services both reduce referral waiting times and with the right technology, improve the evidence available to justify triage decisions. We present and evaluate three different approaches for LLM-based, end-to-end ingestion of variable-length clinical EHR data to assist clinicians when triaging referrals. Our model is able to deliver triage recommendations consistent with existing clinical practices and its architecture was implemented on a single GPU, making it practical for implementation in resource-limited NHS environments where private implementations of LLM technology will be necessary to ensure confidential clinical data are appropriately controlled and governed. Code available at: https://github.com/NtaylorOX/BespokeLLM_Triage.

63A Prospective Comparison of Large Language Models for Early Prediction of Sepsis.PubMed

Supreeth P Shashikumar, Shamim Nemati
Pac Symp Biocomput. 2025;30:109-120. doi: 10.1142/9789819807024_0009.
We present a comparative study on the performance of two popular open-source large language models for early prediction of sepsis: Llama-3 8B and Mixtral 8x7B. The primary goal was to determine whether a smaller model could achieve comparable predictive accuracy to a significantly larger model in the context of sepsis prediction using clinical data.Our proposed LLM-based sepsis prediction system, COMPOSER-LLM, enhances the previously published COMPOSER model, which utilizes structured EHR data to generate hourly sepsis risk scores. The new system incorporates an LLM-based approach to extract sepsis-related clinical signs and symptoms from unstructured clinical notes. For scores falling within high-uncertainty prediction regions, particularly those near the decision threshold, the system uses the LLM to draw additional clinical context from patient notes; thereby enhancing the model's predictive accuracy in challenging diagnostic scenarios.A total of 2,074 patient encounters admitted to the Emergency Department at two hospitals within the University of California San Diego Health system were used for model evaluation in this study. Our findings reveal that the Llama-3 8B model based system (COMPOSER-LLMLlama) achieved a sensitivity of 70.3%, positive predictive value (PPV) of 32.5%, F-1 score of 44.4% and false alarms per patient hour (FAPH) of 0.0194, closely matching the performance of the larger Mixtral 8x7B model based system (COMPOSER-LLMmixtral) which achieved a sensitivity of 72.1%, PPV of 31.9%, F-1 score of 44.2% and FAPH of 0.020. When prospectively evaluated, COMPOSER-LLMLlama demonstrated similar performance to the COMPOSER-LLMmixtral pipeline, with a sensitivity of 68.7%, PPV of 36.6%, F-1 score of 47.7% and FAPH of 0.019 vs. sensitivity of 70.5%, PPV of 36.3%, F-1 score of 47.9% and FAPH of 0.020. This result indicates that, for extraction of clinical signs and symptoms from unstructured clinical notes to enable early prediction of sepsis, the Llama-3 generation of smaller language models can perform as effectively and more efficiently than larger models. This finding has significant implications for healthcare settings with limited resources.

64In-context learning enables multimodal large language models to classify cancer pathology images.PubMed

Dyke Ferber, Georg Wölflein, Isabella C Wiest, et al.
Nat Commun. 2024 Nov 21;15(1):10104. doi: 10.1038/s41467-024-51465-9.
Medical image classification requires labeled, task-specific datasets which are used to train deep learning networks de novo, or to fine-tune foundation models. However, this process is computationally and technically demanding. In language processing, in-context learning provides an alternative, where models learn from within prompts, bypassing the need for parameter updates. Yet, in-context learning remains underexplored in medical image analysis. Here, we systematically evaluate the model Generative Pretrained Transformer 4 with Vision capabilities (GPT-4V) on cancer image processing with in-context learning on three cancer histopathology tasks of high importance: Classification of tissue subtypes in colorectal cancer, colon polyp subtyping and breast tumor detection in lymph node sections. Our results show that in-context learning is sufficient to match or even outperform specialized neural networks trained for particular tasks, while only requiring a minimal number of samples. In summary, this study demonstrates that large vision language models trained on non-domain specific data can be applied out-of-the box to solve medical image-processing tasks in histopathology. This democratizes access of generalist AI models to medical experts without technical background especially for areas where annotated data is scarce.

65Large Foundation Model for Cancer Segmentation.PubMed

Zeyu Ren, Yudong Zhang, Shuihua Wang
Technol Cancer Res Treat. 2024 Jan-Dec;23:15330338241266205. doi: 10.1177/15330338241266205.
Recently, large language models such as ChatGPT have made huge strides in understanding and generating human-like text and have demonstrated considerable success in natural language processing. These foundation models also perform well in computer vision. However, there is a growing need to use these technologies for specific medical tasks, especially for identifying cancer in images. This paper looks at how these foundation models, such as the segment anything model, could be used for cancer segmentation, discussing the potential benefits and challenges of applying large foundation models to help with cancer diagnoses.

66Revolutionizing Digital Pathology With the Power of Generative Artificial Intelligence and Foundation Models.PubMed

Asim Waqas, Marilyn M Bui, Eric F Glassy, et al.
Lab Invest. 2023 Nov;103(11):100255. doi: 10.1016/j.labinv.2023.100255. Epub 2023 Sep 26.
Digital pathology has transformed the traditional pathology practice of analyzing tissue under a microscope into a computer vision workflow. Whole-slide imaging allows pathologists to view and analyze microscopic images on a computer monitor, enabling computational pathology. By leveraging artificial intelligence (AI) and machine learning (ML), computational pathology has emerged as a promising field in recent years. Recently, task-specific AI/ML (eg, convolutional neural networks) has risen to the forefront, achieving above-human performance in many image-processing and computer vision tasks. The performance of task-specific AI/ML models depends on the availability of many annotated training datasets, which presents a rate-limiting factor for AI/ML development in pathology. Task-specific AI/ML models cannot benefit from multimodal data and lack generalization, eg, the AI models often struggle to generalize to new datasets or unseen variations in image acquisition, staining techniques, or tissue types. The 2020s are witnessing the rise of foundation models and generative AI. A foundation model is a large AI model trained using sizable data, which is later adapted (or fine-tuned) to perform different tasks using a modest amount of task-specific annotated data. These AI models provide in-context learning, can self-correct mistakes, and promptly adjust to user feedback. In this review, we provide a brief overview of recent advances in computational pathology enabled by task-specific AI, their challenges and limitations, and then introduce various foundation models. We propose to create a pathology-specific generative AI based on multimodal foundation models and present its potentially transformative role in digital pathology. We describe different use cases, delineating how it could serve as an expert companion of pathologists and help them efficiently and objectively perform routine laboratory tasks, including quantifying image analysis, generating pathology reports, diagnosis, and prognosis. We also outline the potential role that foundation models and generative AI can play in standardizing the pathology laboratory workflow, education, and training.

67Artificial Intelligence in Cardiovascular Disease Prevention: Is it Ready for Prime Time?PubMed

Shyon Parsa, Sulaiman Somani, Ramzi Dudum, et al.
Curr Atheroscler Rep. 2024 Jul;26(7):263-272. doi: 10.1007/s11883-024-01210-w. Epub 2024 May 23.
PURPOSE OF REVIEW: This review evaluates how Artificial Intelligence (AI) enhances atherosclerotic cardiovascular disease (ASCVD) risk assessment, allows for opportunistic screening, and improves adherence to guidelines through the analysis of unstructured clinical data and patient-generated data. Additionally, it discusses strategies for integrating AI into clinical practice in preventive cardiology. RECENT FINDINGS: AI models have shown superior performance in personalized ASCVD risk evaluations compared to traditional risk scores. These models now support automated detection of ASCVD risk markers, including coronary artery calcium (CAC), across various imaging modalities such as dedicated ECG-gated CT scans, chest X-rays, mammograms, coronary angiography, and non-gated chest CT scans. Moreover, large language model (LLM) pipelines are effective in identifying and addressing gaps and disparities in ASCVD preventive care, and can also enhance patient education. AI applications are proving invaluable in preventing and managing ASCVD and are primed for clinical use, provided they are implemented within well-regulated, iterative clinical pathways.

68Radiomics nomogram for predicting chemo-immunotherapy efficiency in advanced non-small cell lung cancer.PubMed

Hua Jin, Yuchao Wang, Xushuo Li, et al.
Sci Rep. 2024 Sep 6;14(1):20788. doi: 10.1038/s41598-024-63415-y.
This study aimed to explore potential radiomics biomarkers in predicting the efficiency of chemo-immunotherapy in patients with advanced non-small cell lung cancer (NSCLC). Eligible patients were prospectively assigned to receive chemo-immunotherapy, and were divided into a primary cohort (n = 138) and an internal validation cohort (n = 58). Additionally, a separative dataset was used as an external validation cohort (n = 60). Radiomics signatures were extracted and selected from the primary tumor sites from chest CT images. A multivariate logistic regression analysis was conducted to identify the independent clinical predictors. Subsequently, a radiomics nomogram model for predicting the efficiency of chemo-immunotherapy was conducted by integrating the selected radiomics signatures and the independent clinical predictors. The receiver operating characteristic (ROC) curves demonstrated that the radiomics model, the clinical model, and the radiomics nomogram model achieved areas under the curve (AUCs) of 0.85 (95% confidence interval [CI] 0.78-0.92), 0.76 (95% CI 0.68-0.84), and 0.89 (95% CI 0.84-0.94), respectively, in the primary cohort. In the internal validation cohort, the corresponding AUCs were 0.93 (95% CI 0.86-1.00), 0.79 (95% CI 0.68-0.91), and 0.96 (95% CI 0.90-1.00) respectively. Moreover, in the external validation cohort, the AUCs were 0.84 (95% CI 0.72-0.96), 0.75 (95% CI 0.62-0.87), and 0.86 (95% CI 0.75-0.96), respectively. In conclusion, the radiomics nomogram provides a convenient model for predicting the effect of chemo-immunotherapy in advanced NSCLC patients.

69CT-based quantification of intratumoral heterogeneity for predicting pathologic complete response to neoadjuvant immunochemotherapy in non-small cell lung cancer.PubMed

Guanchao Ye, Guangyao Wu, Chunyang Zhang, et al.
Front Immunol. 2024 Jun 12;15:1414954. doi: 10.3389/fimmu.2024.1414954. eCollection 2024.
OBJECTIVES: To investigate the prediction of pathologic complete response (pCR) in patients with non-small cell lung cancer (NSCLC) undergoing neoadjuvant immunochemotherapy (NAIC) using quantification of intratumoral heterogeneity from pre-treatment CT image. METHODS: This retrospective study included 178 patients with NSCLC who underwent NAIC at 4 different centers. The training set comprised 108 patients from center A, while the external validation set consisted of 70 patients from center B, center C, and center D. The traditional radiomics model was contrasted using radiomics features. The radiomics features of each pixel within the tumor region of interest (ROI) were extracted. The optimal division of tumor subregions was determined using the K-means unsupervised clustering method. The internal tumor heterogeneity habitat model was developed using the habitats features from each tumor sub-region. The LR algorithm was employed in this study to construct a machine learning prediction model. The diagnostic performance of the model was evaluated using criteria such as area under the receiver operating characteristic curve (AUC), accuracy, specificity, sensitivity, positive predictive value (PPV), and negative predictive value (NPV). RESULTS: In the training cohort, the traditional radiomics model achieved an AUC of 0.778 [95% confidence interval (CI): 0.688-0.868], while the tumor internal heterogeneity habitat model achieved an AUC of 0.861 (95% CI: 0.789-0.932). The tumor internal heterogeneity habitat model exhibits a higher AUC value. It demonstrates an accuracy of 0.815, surpassing the accuracy of 0.685 achieved by traditional radiomics models. In the external validation cohort, the AUC values of the two models were 0.723 (CI: 0.591-0.855) and 0.781 (95% CI: 0.673-0.889), respectively. The habitat model continues to exhibit higher AUC values. In terms of accuracy evaluation, the tumor heterogeneity habitat model outperforms the traditional radiomics model, achieving a score of 0.743 compared to 0.686. CONCLUSION: The quantitative analysis of intratumoral heterogeneity using CT to predict pCR in NSCLC patients undergoing NAIC holds the potential to inform clinical decision-making for resectable NSCLC patients, prevent overtreatment, and enable personalized and precise cancer management.

70Validation and calibration of structural models that combine information from multiple sources.PubMed

Issa J Dahabreh, John B Wong, Thomas A Trikalinos
Expert Rev Pharmacoecon Outcomes Res. 2017 Feb;17(1):27-37. doi: 10.1080/14737167.2017.1277143.
Mathematical models that attempt to capture structural relationships between their components and combine information from multiple sources are increasingly used in medicine. Areas covered: We provide an overview of methods for model validation and calibration and survey studies comparing alternative approaches. Expert commentary: Model validation entails a confrontation of models with data, background knowledge, and other models, and can inform judgments about model credibility. Calibration involves selecting parameter values to improve the agreement of model outputs with data. When the goal of modeling is quantitative inference on the effects of interventions or forecasting, calibration can be viewed as estimation. This view clarifies issues related to parameter identifiability and facilitates formal model validation and the examination of consistency among different sources of information. In contrast, when the goal of modeling is the generation of qualitative insights about the modeled phenomenon, calibration is a rather informal process for selecting inputs that result in model behavior that roughly reproduces select aspects of the modeled phenomenon and cannot be equated to an estimation procedure. Current empirical research on validation and calibration methods consists primarily of methodological appraisals or case-studies of alternative techniques and cannot address the numerous complex and multifaceted methodological decisions that modelers must make. Further research is needed on different approaches for developing and validating complex models that combine evidence from multiple sources.

71Intratumoral and peritumoral radiomics of MRIs predicts pathologic complete response to neoadjuvant chemoimmunotherapy in patients with head and neck squamous cell carcinoma.PubMed

Peiliang Lin, Wenqian Xie, Yong Li, et al.
J Immunother Cancer. 2024 Nov 5;12(11):e009616. doi: 10.1136/jitc-2024-009616.
BACKGROUND: For patients with locally advanced head and neck squamous cell carcinoma (HNSCC), combined programmed death receptor-1 inhibitor and chemotherapy improved response rate to neoadjuvant therapy. However, treatment response varies among patients. There is no tool to predict pathologic complete response (pCR) with high accuracy for now. To develop a tool based on radiomics features of MRI to predict pCR to neoadjuvant chemoimmunotherapy (NACI) may provide valuable assistance in treatment regimen determination for HNSCC. METHODS: From January 2021 to April 2024, a total of 172 patients with HNSCC from three medical center, who received NACI followed by surgery, were included and allocated into a training set (n=84), an internal validation set (n=37) and an external validation set (n=51). Radiomics features were extracted from intratumoral and different peritumoral areas, and radiomics signature (Rad-score) for each area was constructed. A radiomics-clinical nomogram was developed based on Rad-scores and clinicopathological characteristics, tested in the validation sets, and compared with clinical nomogram and combined positive score (CPS) in predicting pCR. RESULTS: The radiomics-clinical nomogram, incorporating peritumoral Rad-score, intratumoral Rad-score and CPS, achieved the highest accuracy with areas under the receiver operating characteristic curve of 0.904 (95% CI, 0.835 to 0.972) in the training cohort, 0.860 (95% CI, 0.722 to 0.998) in the internal validation cohort, and 0.849 (95% CI, 0.739 to 0.959) in the external validation cohort, respectively, which outperformed the clinical nomogram and CPS in predict pCR to NACI for HNSCC. CONCLUSION: A nomogram developed based on intratumoral and peritumoral MRI radiomics features outperformed CPS, a widely employed biomarker, in predict pCR to NACI for HNSCC, which would provide incremental value in treatment regimen determination.

72Prediction models need appropriate internal, internal-external, and external validation.PubMed

Ewout W Steyerberg, Frank E Harrell
J Clin Epidemiol. 2016 Jan;69:245-7. doi: 10.1016/j.jclinepi.2015.04.005. Epub 2015 Apr 18.

73Assessing the performance of QSP models: biology as the driver for validation.PubMed

Fulya Akpinar Singh, Nasrin Afzal, Shepard J Smithline, et al.
J Pharmacokinet Pharmacodyn. 2024 Oct;51(5):533-542. doi: 10.1007/s10928-023-09871-x. Epub 2023 Jun 29.
Validation of a quantitative model is a critical step in establishing confidence in the model's suitability for whatever analysis it was designed. While processes for validation are well-established in the statistical sciences, the field of quantitative systems pharmacology (QSP) has taken a more piecemeal approach to defining and demonstrating validation. Although classical statistical methods can be used in a QSP context, proper validation of a mechanistic systems model requires a more nuanced approach to what precisely is being validated, and what role said validation plays in the larger context of the analysis. In this review, we summarize current thoughts of QSP validation in the scientific community, contrast the aims of statistical validation from several contexts (including inference, pharmacometrics analysis, and machine learning) with the challenges faced in QSP analysis, and use examples from published QSP models to define different stages or levels of validation, any of which may be sufficient depending on the context at hand.

74Development and Validation of a Prediction Model for Thyroid Dysfunction in Patients During Immunotherapy.PubMed

Qian Wang, Tingting Wu, Ru Zhao, et al.
Endocr Pract. 2024 Oct;30(10):943-950. doi: 10.1016/j.eprac.2024.07.006. Epub 2024 Jul 14.
OBJECTIVE: This study was designed to develop and validate a predictive model for assessing the risk of thyroid toxicity following treatment with immune checkpoint inhibitors. METHODS: A retrospective analysis was conducted on a cohort of 586 patients diagnosed with malignant tumors who received programmed cell death 1 (PD-1)/programmed death-ligand 1 (PD-L1) inhibitors. The patients were randomly divided into training and validation cohorts in a 7:3 ratio. Logistic regression analyses were performed on the training set to identify risk factors of thyroid dysfunction, and a nomogram was developed based on these findings. Internal validation was performed using K-fold cross-validation on the validation set. The performance of the nomogram was assessed in terms of discrimination and calibration. Additionally, decision curve analysis was utilized to demonstrate the decision efficiency of the model. RESULTS: Our clinical prediction model consisted of 4 independent predictors of thyroid immune-related adverse events, namely baseline thyrotropin (TSH, OR = 1.427, 95%CI:1.163-1.876), baseline thyroglobulin antibody (TgAb, OR = 1.105, 95%CI:1.035-1.180), baseline thyroid peroxidase antibody (TPOAb, OR = 1.172, 95%CI:1.110-1.237), and baseline platelet count (platelet, OR = 1.004, 95%CI:1.000-1.007). The developed nomogram achieved excellent discrimination with an area under the curve of 0.863 (95%CI: 0.817-0.909) and 0.885 (95%CI: 0.827-0.944) in the training and internal validation cohorts respectively. Calibration curves exhibited a good fit, and the decision curve indicated favorable clinical benefits. CONCLUSION: The proposed nomogram serves as an effective and intuitive tool for predicting the risk of thyroid immune-related adverse events, facilitating clinicians making individualized decisions based on patient-specific information.

75Machine learning in the prediction of immunotherapy response and prognosis of melanoma: a systematic review and meta-analysis.PubMed

Juan Li, Kena Dan, Jun Ai
Front Immunol. 2024 May 21;15:1281940. doi: 10.3389/fimmu.2024.1281940. eCollection 2024.
BACKGROUND: The emergence of immunotherapy has changed the treatment modality for melanoma and prolonged the survival of many patients. However, a handful of patients remain unresponsive to immunotherapy and effective tools for early identification of this patient population are still lacking. Researchers have developed machine learning algorithms for predicting immunotherapy response in melanoma, but their predictive accuracy has been inconsistent. Therefore, the present systematic review and meta-analysis was performed to comprehensively evaluate the predictive accuracy of machine learning in melanoma response to immunotherapy. METHODS: Relevant studies were searched in PubMed, Web of Sciences, Cochrane Library, and Embase from their inception to July 30, 2022. The risk of bias and applicability of the included studies were assessed using the Prediction Model Risk of Bias Assessment Tool (PROBAST). Meta-analysis was performed on R4.2.0. RESULTS: A total of 36 studies consisting of 30 cohort studies and 6 case-control studies were included. These studies were mainly published between 2019 and 2022 and encompassed 75 models. The outcome measures of this study were progression-free survival (PFS), overall survival (OS), and treatment response. The pooled c-index was 0.728 (95%CI: 0.629-0.828) for PFS in the training set, 0.760 (95%CI: 0.728-0.792) and 0.819 (95%CI: 0.757-0.880) for treatment response in the training and validation sets, respectively, and 0.746 (95%CI: 0.721-0.771) and 0.700 (95%CI: 0.677-0.724) for OS in the training and validation sets, respectively. CONCLUSION: Machine learning has considerable predictive accuracy in melanoma immunotherapy response and prognosis, especially in the former. However, due to the lack of external validation and the scarcity of certain types of models, further studies are warranted.

76Development and validation of a new tool to estimate early mortality in patients with advanced cancer treated with immunotherapy.PubMed

Andrea De Giglio, Alessandro Leonetti, Francesca Comito, et al.
Cancer Immunol Immunother. 2024 Oct 3;73(12):246. doi: 10.1007/s00262-024-03836-w.
BACKGROUND: Immune checkpoint inhibitors (ICIs) are standard treatments for advanced solid cancers. Resistance to ICIs, both primary and secondary, poses challenges, with early mortality (EM) within 30-90 days indicating a lack of benefit. Prognostic factors for EM, including the lung immune prognostic index (LIPI), remain underexplored. METHODS: We performed a retrospective, observational study including patients affected by advanced solid tumors, treated with ICI as single agent or combined with other agents. Logistic regression models identified factors associated with EM and 90-day progression risks. A nomogram for predicting 90-day mortality was built and validated within an external cohort. RESULTS: In total, 637 patients received ICIs (single agent or in combination with other drugs) for advanced solid tumors. Most patients were male (61.9%), with NSCLC as the prevalent tumor (61.8%). Within the cohort, 21.3% died within 90 days, 8.4% died within 30 days, and 34.5% experienced early progression. Factors independently associated with 90-day mortality included ECOG PS 2 and a high/intermediate LIPI score. For 30-day mortality, lung metastasis and a high/intermediate LIPI score were independent risk factors. Regarding early progression, high/intermediate LIPI score was independently associated. A predictive nomogram for 90-day mortality combining LIPI and ECOG PS achieved an AUC of 0.76 (95% CI 0.71-0.81). The discrimination ability of the nomogram was confirmed in the external validation cohort (n = 255) (AUC 0.72, 95% CI 0.64-0.80). CONCLUSION: LIPI and ECOG PS independently were able to estimate 90-day mortality, with LIPI also demonstrating prognostic validity for 30-day mortality and early progression.

77Prognostic and predictive value of histogram analysis in patients with non-small cell lung cancer refractory to platinum treated by nivolumab: A multicentre retrospective study.PubMed

Marco Ravanelli, Giorgio Maria Agazzi, Gianluca Milanese, et al.
Eur J Radiol. 2019 Sep;118:251-256. doi: 10.1016/j.ejrad.2019.07.019. Epub 2019 Jul 17.
PURPOSE: The aim of this study was to assess computed-tomography histogram analysis (CTHA) as prognostic and predictive factor in platinum-refractory non-small cell lung carcinoma (NSCLC) treated with immune checkpoint inhibitor Nivolumab. METHOD: One hundred and four patients were enrolled from 3 different centers. CT was performed using similar parameters among different scanners. CTHA was performed with the proprietary software TexRAD, which extracts histogram features at different spatial scale (spatial scale filters, SSF) producing 30 CTHA features per patients. Cross-validated Least Absolute Shrinkage and Selection Operator LASSO was used to select those features which were related to overall and progression-free survival (OS and PFS, respectively). High- and low-risk subgroups were identified using the best cutoff. RESULTS: Median follow-up was 13.8 weeks. Median OS and PFS were 7.3 and 3 months, respectively. LASSO selected kurtosis obtained by SSF = 4 mm as the single feature related to OS, leading to an hazard ratio (HR) of 0.476 (95%CI 0.29-0.77). PFS was related with kurtosis SSF = 6 mm, with HR of 0.556 (95%CI 0.36-0.86). CONCLUSION: Despite its limitations, this study is the first which suggests that CTHA could play a role in stratifying prognosis and treatment response in patients with NSCLC treated with Nivolumab.

78dsOMOP: bridging OMOP CDM and DataSHIELD for secure federated analysis of standardized clinical data.PubMed

David Sarrat-González, Xavier Escribà-Montagut, Jared Houghtaling, et al.
Bioinformatics. 2025 Jun 2;41(6). doi: 10.1093/bioinformatics/btaf286.
MOTIVATION: Collaborative clinical research projects face several challenges related to data sharing. The disparity between data standards and strict privacy regulations become more relevant as the number of involved institutions increases. To address these challenges, the scientific community has progressively adopted common data models like the Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) for multicenter data standardization and implemented federated data analysis platforms like DataSHIELD to perform remote analyses without transferring individual-level data between centers, thus mitigating disclosure risks. However, there is no native implementation that automatically combines both solutions, revealing the need for a tool that enables interoperability between these systems. RESULTS: We present dsOMOP, a collection of DataSHIELD packages that facilitates automated extraction and transformation of OMOP CDM data into DataSHIELD-compatible datasets, enabling disclosure-controlled federated analyses of standardized clinical data. dsOMOP allows research institutions to provide access to their data for collaborative projects in a format that is interoperable with the project's available data, thus facilitating the analysis of large-scale, multicenter clinical data. It incorporates OMOP data directly into the DataSHIELD workflow, where all analyses occur entirely in a federated environment subject to rigorous disclosure controls, ensuring that only aggregated, non-disclosive results are ever returned to analysts. AVAILABILITY AND IMPLEMENTATION: The general information page for the dsOMOP environment is available at https://isglobal-brge.github.io/dsOMOP, where the most recent installation instructions and usage guides for all dsOMOP packages and their extensions can be found in the "Packages" section.The dsOMOP package and its complementary tools are fully available under the MIT license on GitHub: dsOMOP (https://github.com/isglobal-brge/dsOMOP), dsOMOPClient (https://github.com/isglobal-brge/dsOMOPClient), dsOMOPHelper (https://github.com/isglobal-brge/dsOMOPHelper), and dsOMOP.oracle (https://github.com/isglobal-brge/dsOMOP.oracle).Usage vignettes for the client-side packages are available at the websites of dsOMOPClient (https://isglobal-brge.github.io/dsOMOPClient) and dsOMOPHelper (https://isglobal-brge.github.io/dsOMOPHelper). A permanent archival snapshot of the exact code used in this manuscript is deposited at Figshare: https://doi.org/10.6084/m9.figshare.28607186.

79POPCORN: A web service for individual PrognOsis prediction based on multi-center clinical data CollabORatioN without patient-level data sharing.PubMed

Yu Tian, Yong Shang, Dan-Yang Tong, et al.
J Biomed Inform. 2018 Oct;86:1-14. doi: 10.1016/j.jbi.2018.08.008. Epub 2018 Aug 10.
BACKGROUND AND OBJECTIVE: Clinical prognosis prediction plays an important role in clinical research and practice. The construction of prediction models based on electronic health record data has recently become a research focus. Due to the lack of external validation, prediction models based on single-center, hospital-specific datasets may not perform well with datasets from other medical institutions. Therefore, research investigating prognosis prediction model construction based on a collaborative analysis of multi-center electronic health record data could increase the number and coverage of patients used for model training, enrich patient prognostic features and ultimately improve the accuracy and generalization of prognosis prediction. MATERIALS AND METHODS: A web service for individual prognosis prediction based on multi-center clinical data collaboration without patient-level data sharing (POPCORN) was proposed. POPCORN focuses on solving key issues in multi-center collaborative research based on electronic health record systems; these issues include the standardization of clinical data expression, the preservation of patient privacy during model training and the effect of case mix variance on the prediction model construction and application. POPCORN is based on a multivariable meta-analysis and a Bayesian framework and can construct suitable prediction models for multiple clinical scenarios that can effectively adapt to complex clinical application environments. RESULTS: POPCORN was validated using a joint, multi-center collaborative research network between China and the United States with patients diagnosed with colorectal cancer. The performance of the models based on POPCORN was comparable to that of the standard prognosis prediction model; however, POPCORN did not expose raw patient data. The prediction models had similar AUC, but the BMA model had the lowest ECI across all prediction models, indicating that this model had better calibration performance than the other models, especially for patients in Chinese hospitals. CONCLUSIONS: The POPCORN system can build prediction models that perform well in complex clinical application scenarios and can provide effective decision support for individual patient prognostic predictions.

80The false hope of current approaches to explainable artificial intelligence in health care.PubMed

Marzyeh Ghassemi, Luke Oakden-Rayner, Andrew L Beam
Lancet Digit Health. 2021 Nov;3(11):e745-e750. doi: 10.1016/S2589-7500(21)00208-9.
The black-box nature of current artificial intelligence (AI) has caused some to question whether AI must be explainable to be used in high-stakes scenarios such as medicine. It has been argued that explainable AI will engender trust with the health-care workforce, provide transparency into the AI decision making process, and potentially mitigate various kinds of bias. In this Viewpoint, we argue that this argument represents a false hope for explainable AI and that current explainability methods are unlikely to achieve these goals for patient-level decision support. We provide an overview of current explainability techniques and highlight how various failure cases can cause problems for decision making for individual patients. In the absence of suitable explainability methods, we advocate for rigorous internal and external validation of AI models as a more direct means of achieving the goals often associated with explainability, and we caution against having explainability be a requirement for clinically deployed models.

81Barriers and facilitators perceived by physicians when using prediction models in practice.PubMed

Teus H Kappen, Kim van Loon, Martinus A M Kappen, et al.
J Clin Epidemiol. 2016 Feb;70:136-45. doi: 10.1016/j.jclinepi.2015.09.008. Epub 2015 Sep 21.
OBJECTIVES: Prediction models may facilitate risk-based management of health care conditions. In a large cluster-randomized trial, presenting calculated risks of postoperative nausea and vomiting (PONV) to physicians (assistive approach) increased risk-based management of PONV. This increase did not improve patient outcome-that is, PONV incidence. This prompted us to explore how prediction tools guide the decision-making process of physicians. STUDY DESIGN AND SETTING: Using mixed methods, we interviewed eight physicians to understand how predicted risks were perceived by the physicians and how they influenced decision making. Subsequently, all 57 physicians of the trial were surveyed for how the presented risks influenced their perceptions. RESULTS: Although the prediction tool made physicians more aware of PONV prevention, the physicians reported three barriers to use predicted risks in their decision making. PONV was not considered an outcome of utmost importance; decision making on PONV prophylaxis was mostly intuitive rather than risk based; prediction models do not weigh benefits and risks of prophylactic drugs. CONCLUSION: Combining probabilistic output of the model with their clinical experience may be difficult for physicians, especially when their decision-making process is mostly intuitive. Adding recommendations to predicted risks (directive approach) was considered an important step to facilitate the uptake of a prediction tool.

82A governance model for the application of AI in health care.PubMed

Sandeep Reddy, Sonia Allan, Simon Coghlan, et al.
J Am Med Inform Assoc. 2020 Mar 1;27(3):491-497. doi: 10.1093/jamia/ocz192.
As the efficacy of artificial intelligence (AI) in improving aspects of healthcare delivery is increasingly becoming evident, it becomes likely that AI will be incorporated in routine clinical care in the near future. This promise has led to growing focus and investment in AI medical applications both from governmental organizations and technological companies. However, concern has been expressed about the ethical and regulatory aspects of the application of AI in health care. These concerns include the possibility of biases, lack of transparency with certain AI algorithms, privacy concerns with the data used for training AI models, and safety and liability issues with AI application in clinical environments. While there has been extensive discussion about the ethics of AI in health care, there has been little dialogue or recommendations as to how to practically address these concerns in health care. In this article, we propose a governance model that aims to not only address the ethical and regulatory issues that arise out of the application of AI in health care, but also stimulate further discussion about governance of AI in health care.

83[Artificial intelligence in medicine: limits and obstacles.].PubMed

Eugenio Santoro
Recenti Prog Med. 2017 Dec;108(12):500-502. doi: 10.1701/2829.28580.
Data scientists and physicians are starting to use artificial intelligence (AI) even in the medical field in order to better understand the relationships among the huge amount of data coming from the great number of sources today available. Through the data interpretation methods made available by the recent AI tools, researchers and AI companies have focused on the development of models allowing to predict the risk of suffering from a specific disease, to make a diagnosis, and to recommend a treatment that is based on the best and most updated scientific evidence. Even if AI is used to perform unimaginable tasks until a few years ago, the awareness about the ongoing revolution has not yet spread through the medical community for several reasons including the lack of evidence about safety, reliability and effectiveness of these tools, the lack of regulation accompanying hospitals in the use of AI by health care providers, the difficult attribution of liability in case of errors and malfunctions of these systems, and the ethical and privacy questions that they raise and that, as of today, are still unanswered.