通过相互匹配关系建模增强医学视觉语言对比学习

Enhancing Medical Vision-Language Contrastive Learning via Inter-Matching Relation Modeling.

作者信息

Li Mingjian, Meng Mingyuan, Fulham Michael, Feng David Dagan, Bi Lei, Kim Jinman

出版信息

IEEE Trans Med Imaging. 2025 Jun;44(6):2463-2476. doi: 10.1109/TMI.2025.3534436.

DOI:10.1109/TMI.2025.3534436

Abstract

Medical image representations can be learned through medical vision-language contrastive learning (mVLCL) where medical imaging reports are used as weak supervision through image-text alignment. These learned image representations can be transferred to and benefit various downstream medical vision tasks such as disease classification and segmentation. Recent mVLCL methods attempt to align image sub-regions and the report keywords as local-matchings. However, these methods aggregate all local-matchings via simple pooling operations while ignoring the inherent relations between them. These methods therefore fail to reason between local-matchings that are semantically related, e.g., local-matchings that correspond to the disease word and the location word (semantic-relations), and also fail to differentiate such clinically important local-matchings from others that correspond to less meaningful words, e.g., conjunction words (importance-relations). Hence, we propose a mVLCL method that models the inter-matching relations between local-matchings via a relation-enhanced contrastive learning framework (RECLF). In RECLF, we introduce a semantic-relation reasoning module (SRM) and an importance-relation reasoning module (IRM) to enable more fine-grained report supervision for image representation learning. We evaluated our method using six public benchmark datasets on four downstream tasks, including segmentation, zero-shot classification, linear classification, and cross-modal retrieval. Our results demonstrated the superiority of our RECLF over the state-of-the-art mVLCL methods with consistent improvements across single-modal and cross-modal tasks. These results suggest that our RECLF, by modeling the inter-matching relations, can learn improved medical image representations with better generalization capabilities.

摘要

医学图像表示可以通过医学视觉-语言对比学习（mVLCL）来学习，其中医学成像报告通过图像-文本对齐用作弱监督。这些学习到的图像表示可以迁移到各种下游医学视觉任务（如疾病分类和分割）并使其受益。最近的mVLCL方法试图将图像子区域和报告关键词对齐为局部匹配。然而，这些方法通过简单的池化操作聚合所有局部匹配，同时忽略它们之间的内在关系。因此，这些方法无法在语义相关的局部匹配之间进行推理，例如与疾病词和位置词对应的局部匹配（语义关系），也无法将这种临床上重要的局部匹配与其他对应于意义较小的词（如连词）的局部匹配区分开来（重要性关系）。因此，我们提出了一种mVLCL方法，该方法通过关系增强对比学习框架（RECLF）对局部匹配之间的匹配关系进行建模。在RECLF中，我们引入了语义关系推理模块（SRM）和重要性关系推理模块（IRM），以实现对图像表示学习更细粒度的报告监督。我们在四个下游任务（包括分割、零样本分类、线性分类和跨模态检索）上使用六个公共基准数据集对我们的方法进行了评估。我们的结果表明，我们的RECLF优于现有最先进的mVLCL方法，在单模态和跨模态任务中都有持续的改进。这些结果表明，我们的RECLF通过对匹配关系进行建模，可以学习到具有更好泛化能力的改进医学图像表示。

相似文献

Enhancing Medical Vision-Language Contrastive Learning via Inter-Matching Relation Modeling.通过相互匹配关系建模增强医学视觉语言对比学习

IEEE Trans Med Imaging. 2025 Jun;44(6):2463-2476. doi: 10.1109/TMI.2025.3534436.

Short-Term Memory Impairment短期记忆障碍

Boundary-aware information maximization for self-supervised medical image segmentation.用于自监督医学图像分割的边界感知信息最大化

Med Image Anal. 2024 May;94:103150. doi: 10.1016/j.media.2024.103150. Epub 2024 Mar 28.

Label-Free Medical Image Quality Evaluation by Semantics-Aware Contrastive Learning in IoMT.物联网环境下基于语义感知对比学习的无标签医学图像质量评估

IEEE J Biomed Health Inform. 2025 Apr;29(4):2335-2344. doi: 10.1109/JBHI.2023.3340201. Epub 2025 Apr 4.

CACL: Cluster-Aware Adversarial Contrastive Learning for Pathological Image Analysis.CACL：用于病理图像分析的聚类感知对抗对比学习

IEEE J Biomed Health Inform. 2025 Jul;29(7):5095-5108. doi: 10.1109/JBHI.2025.3552640.

Comparison of Two Modern Survival Prediction Tools, SORG-MLA and METSSS, in Patients With Symptomatic Long-bone Metastases Who Underwent Local Treatment With Surgery Followed by Radiotherapy and With Radiotherapy Alone.两种现代生存预测工具 SORG-MLA 和 METSSS 在接受手术联合放疗和单纯放疗治疗有症状长骨转移患者中的比较。

Clin Orthop Relat Res. 2024 Dec 1;482(12):2193-2208. doi: 10.1097/CORR.0000000000003185. Epub 2024 Jul 23.

Sexual Harassment and Prevention Training性骚扰与预防培训

A New Measure of Quantified Social Health Is Associated With Levels of Discomfort, Capability, and Mental and General Health Among Patients Seeking Musculoskeletal Specialty Care.一种新的量化社会健康指标与寻求肌肉骨骼专科护理的患者的不适程度、能力以及心理和总体健康水平相关。

Clin Orthop Relat Res. 2025 Apr 1;483(4):647-663. doi: 10.1097/CORR.0000000000003394. Epub 2025 Feb 5.

Are Current Survival Prediction Tools Useful When Treating Subsequent Skeletal-related Events From Bone Metastases?当前的生存预测工具在治疗骨转移后的骨骼相关事件时有用吗？

Clin Orthop Relat Res. 2024 Sep 1;482(9):1710-1721. doi: 10.1097/CORR.0000000000003030. Epub 2024 Mar 22.

Exploring the Potential of Electroencephalography Signal-Based Image Generation Using Diffusion Models: Integrative Framework Combining Mixed Methods and Multimodal Analysis.利用扩散模型探索基于脑电图信号的图像生成潜力：结合混合方法和多模态分析的综合框架

JMIR Med Inform. 2025 Jun 25;13:e72027. doi: 10.2196/72027.

引用本文的文献

MedAlmighty: enhancing disease diagnosis with large vision model distillation.MedAlmighty：利用大型视觉模型蒸馏技术提升疾病诊断能力。

Front Artif Intell. 2025 Aug 12;8:1527980. doi: 10.3389/frai.2025.1527980. eCollection 2025.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

通过相互匹配关系建模增强医学视觉语言对比学习

Enhancing Medical Vision-Language Contrastive Learning via Inter-Matching Relation Modeling.

作者信息

出版信息

相似文献

引用本文的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

引用本文的文献