• Suppr超能文献
  • 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
定价套餐&价格
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Win 客户端微信小程序
定价
会员套餐积分包API 积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026
  1. 首页
  2. 分享广场
  3. 手术视频大语言模型:进展、应用与未来挑战

手术视频大语言模型:进展、应用与未来挑战

文献检索匿名用户发表于 2026年05月03日 12:047阅读
检索问题

手术视频大语言模型训练,最新进展和应用前景

在手术视频大语言模型(LLM)训练领域,最新的进展主要集中在解决数据稀疏性、异质性、隐私问题以及如何更有效地利用多模态数据和专家知识。这些进展极大地拓展了LLM在手术领域的应用前景,从辅助教育和培训到提升手术安全性和效率,再到自动化文档和决策支持。

最新进展

  1. 多模态LLM与自适应学习框架

    • LMT++框架:一项研究提出了LMT++,一个LLM辅助的多教师持续学习(CL)框架,用于解决机器人手术视频中的持续视觉问答(VQA)问题 。该框架利用多模态LLM作为补充教师,其强大的泛化能力和对手术领域的理解能够弥补因不同来源手术程序引起的显著领域偏移和医疗器械或手术程序出现频率不均导致的数据不平衡问题 。
    • 创新的数据处理和权重分配:为了将LLM整合到CL框架中,研究提出了一种创新方法,将复杂的LLM嵌入转换为CL训练框架中使用的logits值 。此外,还设计了一种自适应权重分配方法,以平衡LLM的泛化能力和传统VQA模型在先前模型训练过程中获得的领域专业知识 。
    • 新的数据集与超越SOTA:研究还创建了一个新的手术VQA数据集用于模型评估,实验结果表明LMT++超越了现有最先进的CL方法 。
  2. 大规模、分层的手术视觉语言数据集 (SurgLaVi)

    • 解决数据限制:现有的外科VLP(Vision-Language Pre-training)在规模、程序多样性、语义质量和分层结构方面受到限制,阻碍了手术VLP的进展 。
    • SurgLaVi的构建:为解决这些问题,SurgLaVi被提出,它是迄今为止最大、最多样化的手术视觉语言数据集,包含近24万个片段-字幕对,涵盖200多种手术程序,并具有粗粒度、中粒度和细粒度的分层级别 。其核心是一个全自动流程,系统地生成手术视频的细粒度转录,并将其分割成连贯的程序单元,通过双模态过滤去除不相关和噪声样本,以确保高质量的注释 。
    • SurgLaVi-β的发布:为了提高可访问性,还发布了SurgLaVi-β,一个包含11.3万个片段-字幕对的开源衍生数据集,其规模是现有外科VLP数据集的四倍多 。
    • SurgCLIP基准模型:研究引入了SurgCLIP,一个CLIP风格的视频-文本对比框架,作为代表性的基础模型,通过在阶段、步骤、动作和工具识别方面实现显著改进,超越了现有最先进的方法 。这验证了大规模、语义丰富和分层结构的数据集能够直接转化为更强大和更具泛化性的表征,使SurgLaVi成为开发手术基础模型的关键资源 。
  3. 利用专家知识和语言进行手术视频理解

    • 模仿人类学习过程:一项研究提出了一种新颖的方法来解决标注训练数据的稀疏性和异质性问题,其灵感来源于人类观察专家并理解其解释的学习过程 。
    • 视频-语言模型的训练:该方法利用一个在对齐、去噪和生成任务上训练的视频-语言模型,学习短时空和多模态表示 。随后,一个任务特定的时间模型被用于捕捉整个视频中的关系 。
    • 大规模预训练数据集的构建:为了实现手术领域的全面视频-语言理解,研究引入了一种数据收集和过滤策略,从教育性YouTube视频构建了一个大规模预训练数据集 。
    • 参数高效微调:通过将公开可用的手术数据集中的下游任务注释投射到语言领域,实现了参数高效的微调 。
    • 性能提升与新任务能力:该方法在两项手术领域的实验中表现出高达7%的阶段分割任务性能提升,5%的零样本阶段分割能力,并在少量样本设置中与完全监督模型具有可比性 。该模型还首次为手术视频提供了密集视频字幕(DVC)的综合解决方案,尽管手术领域缺乏现有的DVC数据集 。
  4. 文本到视频生成AI (Sora)

    • 创新技术:Sora是生成式AI领域的创新,它利用自然语言处理、深度学习和计算机视觉从文本提示生成令人印象深刻的视频 。
    • 潜在应用:在神经外科领域,Sora在患者教育、公共卫生、手术培训和规划以及研究传播方面具有许多潜在应用 。

应用前景

LLM在手术视频训练中的进展预示着其在多个方面具有广泛的应用前景,旨在提高手术的安全性、效率和教学质量。

  1. 手术培训与教育

    • 个性化反馈与技能发展:LLM可以增强沟通、个性化反馈并促进技能发展,尤其在基于模拟的训练、AI驱动的评估工具、视频评估系统以及虚拟现实(VR)和增强现实(AR)平台中发挥作用 。它们还可以用于反馈的转录、翻译和总结 。
    • 辅助外科医生:通过对术中视频的分析,LLM可以帮助识别手术阶段、器械使用,甚至预测手术持续时间,从而为外科医生提供实时辅助和教育 。这使得AI模型能够更好地理解手术过程,并为外科医生提供更安全、更有效的护理 。
    • 精细化活动识别:识别手术动作的三联体(器械、动词、目标)组合可以提供关于手术活动更全面的细节,这对于开发手术中的AI辅助至关重要 。虽然目前手术工作流分析尚未完全解决,但这些进展为未来细粒度手术活动识别的研究指明了方向 。
    • 生成式视频用于教学:Sora等文本到视频的生成式AI工具,可以从文本提示生成高质量的手术视频,用于患者教育、公共卫生宣导和手术培训与规划,使得复杂的医疗信息更易于理解和传播 。
  2. 手术规划与决策支持

    • 术中视频分析:计算机视觉和AI的快速发展,特别是深度学习的应用,使得计算机能够准确理解图像并记住过去的手术事件,从而能够识别各种手术中的手术阶段和器械 。未来研究将利用更大的视频数据集和改进的算法,提高准确性和可解释性,创建有临床实用价值的AI模型,从而增强外科医生提供更安全护理的能力 。
    • 风险预测与决策辅助:基于机器学习的算法正被研究用于风险预测的决策辅助,甚至用于术中应用,包括图像识别和视频分析 。
    • 工作流理解:通过对齐语言与手术视频,视觉语言预训练(VLP)可以实现工作流理解和跨任务的知识转移,而无需依赖专家标注的数据集 。
  3. 自动化文档与报告

    • 提高操作报告的准确性:AI计算机视觉算法能够自动检测手术步骤,并将每个检测到的步骤映射到预设文本,从而编译成叙述性的AI操作报告 。研究表明,AI创建的操作报告比外科医生手写的报告具有更高的准确性(87.3% vs 72.8%),并且临床显著差异更少(12.7% vs 27.2%) 。
    • 减轻行政负担:自动创建视频驱动的AI手术操作报告有望减少文档负担,提高操作报告的准确性,促进手术透明度,并减少手术文档的主观性,从而潜在地缓解外科医生倦怠 。
  4. 提高手术安全性和效率

    • 上下文感知决策支持:通过对手术工作流的实时反馈分析,可以实现手术室中的上下文感知决策支持,从而提高手术安全性和效率 。
    • 辅助术中决策:AI在手术室中的现有应用包括手术持续时间预测、手势识别、术中癌症检测、术中视频分析、工作流识别、内窥镜引导系统、打结和骨科手术中骨骼的自动注册和跟踪等,这些都有望提高手术精度,减少人力,支持术中决策,并提高手术安全性 。
    • 克服挑战:尽管AI带来了巨大的机遇,但仍需解决一些挑战,包括准确性和可靠性、伦理和隐私问题、AI模型中的偏见、与现有培训系统的整合以及AI辅助工具的培训和采用 。通过积极应对这些挑战,未来手术培训的格局可能会被重塑,为受训者提供更全面、安全和有效的学习体验,最终带来更好的患者结局 。

挑战与未来方向

尽管LLM在手术视频训练中展现出巨大潜力,但仍存在一些挑战需要克服:

  • 数据隐私和伦理:患者数据的隐私问题是一个核心挑战,这使得训练VQA模型时使用以前的数据受到限制,从而需要采用无样本持续学习方法 。此外,偏见、伦理和隐私问题是AI模型集成到手术领域必须解决的典型担忧 。
  • 模型准确性和可靠性:虽然AI在某些任务中表现出更高的准确性,但物理上不合理的动作生成、主体自发生成、物体不自然变形、不准确的物理交互以及多主体生成时异常行为表现等是当前模型面临的显著局限性 。需要进一步调查以确定生成视频与非生成视频相比的有效性 。
  • 整合与采纳:将AI工具整合到现有培训系统和临床实践中,以及医生和培训生的培训和采纳,都是实现广泛应用的关键 。
  • 临床验证:目前AI在手术室中的实际应用仍然有限,未来的研究应侧重于AI应用的临床验证、法律和伦理考量以及实施轨迹的评估 。
  • 基础模型开发:SurgLaVi等大规模数据集的出现,为开发更强大、更具泛化性的手术基础模型奠定了基础,但如何有效地利用这些数据构建能够理解复杂手术工作流、进行精细化交互分析的通用模型,仍是重要的研究方向 。
  • 多模态融合的深度:虽然多模态LLM和视频-语言模型已经取得进展,但如何更深层次地融合视觉、语言和结构化医疗数据,以实现更细致的上下文理解和推理,是进一步提升模型性能的关键 。

总而言之,手术视频大语言模型的训练正处于快速发展阶段,通过解决数据限制、融合多模态信息和利用先进的深度学习技术,其在手术培训、规划、文档和决策支持等方面的应用前景广阔。然而,要实现这些潜力的全面释放,还需要在数据隐私、模型可靠性、伦理规范和临床验证方面进行持续的努力。

References

1Text-to-video generative artificial intelligence: sora in neurosurgery.PubMed

Ali A Mohamed, Brandon Lucke-Wold
Artificial intelligence (AI) has increased in popularity in neurosurgery, with recent interest in generative AI algorithms such as the Large Language Model (LLM) ChatGPT. Sora, an innovation in generative AI, leverages natural language processing, deep learning, and computer vision to generate impressive videos from text prompts. This new tool has many potential applications in neurosurgery. These include patient education, public health, surgical training and planning, and research dissemination. However, there are considerable limitations to the current model such as physically implausible motion generation, spontaneous generation of subjects, unnatural object morphing, inaccurate physical interactions, and abnormal behavior presentation when many subjects are generated. Other typical concerns are with respect to patient privacy, bias, and ethics. Further, appropriate investigation is required to determine how effective generative videos are compared to their non-generated counterparts, irrespective of any limitations. Despite these challenges, Sora and other iterations of its text-to-video generative application may have many benefits to the neurosurgical community.

2LMT++: Adaptively Collaborating LLMs With Multi-Specialized Teachers for Continual VQA in Robotic Surgical Videos.PubMed

Yuyang Du, Kexin Chen, Yue Zhan, et al.
Visual question answering (VQA) plays a vital role in advancing surgical education. However, due to the privacy concern of patient data, training VQA model with previously used data becomes restricted, making it necessary to use the exemplar-free continual learning (CL) approach. Previous CL studies in the surgical field neglected two critical issues: i) significant domain shifts caused by the wide range of surgical procedures collected from various sources, and ii) the data imbalance problem caused by the unequal occurrence of medical instruments or surgical procedures. This paper addresses these challenges with a multimodal large language model (LLM) and an adaptive weight assignment strategy. First, we developed a novel LLM-assisted multi-teacher CL framework (named LMT++), which could harness the strength of a multimodal LLM as a supplementary teacher. The LLM's strong generalization ability, as well as its good understanding of the surgical domain, help to address the knowledge gap arising from domain shifts and data imbalances. To incorporate the LLM in our CL framework, we further proposed an innovative approach to process the training data, which involves the conversion of complex LLM embeddings into logits value used within our CL training framework. Moreover, we design an adaptive weight assignment approach that balances the generalization ability of the LLM and the domain expertise of conventional VQA models obtained in previous model training processes within the CL framework. Finally, we created a new surgical VQA dataset for model evaluation. Comprehensive experimental findings on these datasets show that our approach surpasses state-of-the-art CL methods.

3Computer vision in surgery.PubMed

Thomas M Ward, Pietro Mascagni, Yutong Ban, et al.
The fields of computer vision (CV) and artificial intelligence (AI) have undergone rapid advancements in the past decade, many of which have been applied to the analysis of intraoperative video. These advances are driven by wide-spread application of deep learning, which leverages multiple layers of neural networks to teach computers complex tasks. Prior to these advances, applications of AI in the operating room were limited by our relative inability to train computers to accurately understand images with traditional machine learning (ML) techniques. The development and refining of deep neural networks that can now accurately identify objects in images and remember past surgical events has sparked a surge in the applications of CV to analyze intraoperative video and has allowed for the accurate identification of surgical phases (steps) and instruments across a variety of procedures. In some cases, CV can even identify operative phases with accuracy similar to surgeons. Future research will likely expand on this foundation of surgical knowledge using larger video datasets and improved algorithms with greater accuracy and interpretability to create clinically useful AI models that gain widespread adoption and augment the surgeon's ability to provide safer care for patients everywhere.

4Innovations in surgical training: exploring the role of artificial intelligence and large language models (LLM).PubMed

Julian Varas, Brandon Valencia Coronel, Ignacio Villagrán, et al.
The landscape of surgical training is rapidly evolving with the advent of artificial intelligence (AI) and its integration into education and simulation. This manuscript aims to explore the potential applications and benefits of AI-assisted surgical training, particularly the use of large language models (LLMs), in enhancing communication, personalizing feedback, and promoting skill development. We discuss the advancements in simulation-based training, AI-driven assessment tools, video-based assessment systems, virtual reality (VR) and augmented reality (AR) platforms, and the potential role of LLMs in the transcription, translation, and summarization of feedback. Despite the promising opportunities presented by AI integration, several challenges must be addressed, including accuracy and reliability, ethical and privacy concerns, bias in AI models, integration with existing training systems, and training and adoption of AI-assisted tools. By proactively addressing these challenges and harnessing the potential of AI, the future of surgical training may be reshaped to provide a more comprehensive, safe, and effective learning experience for trainees, ultimately leading to better patient outcomes. .

5Enhancing Accuracy of Operative Reports with Automated Artificial Intelligence Analysis of Surgical Video.PubMed

Abhinav Khanna, Tamir Wolf, Igor Frank, et al.
BACKGROUND: The creation of operative reports is a tedious documentation task that increases administrative burden, which is a potential driver of burnout. Additionally, operative reports are inherently subjective and may contain inaccuracies and incomplete information. Recent advances in artificial intelligence (AI) technology have enabled computer vision systems to accurately detect operative steps on surgical video. We aim to develop a platform for the automated creation of video-based AI surgical operative reports in robotic-assisted radical prostatectomy. STUDY DESIGN: An AI computer vision algorithm was used to automatically detect surgical steps. Each step detected on video was mapped to prespecified text, which was then compiled into a narrative AI operative report. The accuracy of the AI operative reports was compared with operative reports written by surgeons, using expert review of raw surgical video footage as the ground truth. RESULTS: A total of 158 cases from a tertiary referral center were included. In surgeon operative reports, 84 cases (53.2%) had at least 1 discrepancy between the operative report and surgical video. Of these, 43 cases (27.2%) had a clinically significant discrepancy based on expert video review. In AI operative reports, 46 cases (29.1%) had at least 1 discrepancy between the operative report and surgical video. Of these, 20 cases (12.7%) had a clinically significant discrepancy. Overall accuracy was higher for AI operative reports compared with surgeon operative reports (87.3% vs 72.8%, p = 0.001). CONCLUSIONS: Operative reports created by AI achieved higher accuracy than those written by surgeons. To our knowledge, this is the first report of automated video-based AI surgical documentation. Future studies are warranted to explore the potential for this novel tool to reduce documentation burden, improve operative report accuracy, promote surgical transparency, and decrease subjectivity in surgical documentation.

6SurgLaVi: Large-scale hierarchical dataset for surgical vision-language representation learning.PubMed

Alejandra Perez, Chinedu Nwoye, Ramtin Raji Kermani, et al.
Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in surgical VLP remains constrained by the limited scale, procedural diversity, semantic quality, and hierarchical structure of existing datasets. In this work, we present SurgLaVi, the largest and most diverse surgical vision-language dataset to date, comprising nearly 240k clip-caption pairs from more than 200 procedures, and featuring hierarchical levels at coarse-, mid-, and fine-level. At the core of SurgLaVi lies a fully automated pipeline that systematically generates fine-grained transcriptions of surgical videos and segments them into coherent procedural units. To ensure high-quality annotations, it applies dual-modality filtering to remove irrelevant and noisy samples. Within this framework, the resulting captions are enriched with contextual detail, producing annotations that are both semantically rich and easy to interpret. To ensure accessibility, we release SurgLaVi-β, an open-source derivative of 113k clip-caption pairs constructed entirely from public data, which is over four times larger than existing surgical VLP datasets. To demonstrate the value of the SurgLaVi datasets, we introduce SurgCLIP, a CLIP-style video-text contrastive framework with dual encoders, as a representative base model. SurgCLIP achieves consistent improvements across phase, step, action, and tool recognition, surpassing prior state-of-the-art methods, often by large margins. These results validate that large-scale, semantically rich, and hierarchically structured datasets directly translate into stronger and more generalizable representations, establishing SurgLaVi as a key resource for developing surgical foundation models.

7CholecTriplet2021: A benchmark challenge for surgical action triplet recognition.PubMed

Chinedu Innocent Nwoye, Deepak Alapatt, Tong Yu, et al.
Context-aware decision support in the operating room can foster surgical safety and efficiency by leveraging real-time feedback from surgical workflow analysis. Most existing works recognize surgical activities at a coarse-grained level, such as phases, steps or events, leaving out fine-grained interaction details about the surgical activity; yet those are needed for more helpful AI assistance in the operating room. Recognizing surgical actions as triplets of ‹instrument, verb, target› combination delivers more comprehensive details about the activities taking place in surgical videos. This paper presents CholecTriplet2021: an endoscopic vision challenge organized at MICCAI 2021 for the recognition of surgical action triplets in laparoscopic videos. The challenge granted private access to the large-scale CholecT50 dataset, which is annotated with action triplet information. In this paper, we present the challenge setup and the assessment of the state-of-the-art deep learning methods proposed by the participants during the challenge. A total of 4 baseline methods from the challenge organizers and 19 new deep learning algorithms from the competing teams are presented to recognize surgical action triplets directly from surgical videos, achieving mean average precision (mAP) ranging from 4.2% to 38.1%. This study also analyzes the significance of the results obtained by the presented approaches, performs a thorough methodological comparison between them, in-depth result analysis, and proposes a novel ensemble method for enhanced recognition. Our analysis shows that surgical workflow analysis is not yet solved, and also highlights interesting directions for future research on fine-grained surgical activity recognition which is of utmost importance for the development of AI in surgery.

8Watch and learn: leveraging expert knowledge and language for surgical video understanding.PubMed

David Gastager, Ghazal Ghazaei, Constantin Patsch
PURPOSE: Automated surgical workflow analysis is a common yet challenging task with diverse applications in surgical education, research, and clinical decision-making. Although videos are commonly collected during surgical interventions, the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions. We introduce a novel approach for addressing the sparsity and heterogeneity of annotated training data inspired by the human learning procedure of watching experts and understanding their explanations. METHODS: Our method leverages a video-language model trained on alignment, denoising, and generative tasks to learn short-term spatio-temporal and multimodal representations. A task-specific temporal model is then used to capture relationships across entire videos. To achieve comprehensive video-language understanding in the surgical domain, we introduce a data collection and filtering strategy to construct a large-scale pretraining dataset from educational YouTube videos. We then utilize parameter-efficient fine-tuning by projecting downstream task annotations from publicly available surgical datasets into the language domain. RESULTS: Extensive experiments in two surgical domains demonstrate the effectiveness of our approach, with performance improvements of up to 7% in phase segmentation tasks, 5% in zero-shot phase segmentation, and comparable capabilities to fully supervised models in few-shot settings. Harnessing our model's capabilities for long-range temporal localization and text generation, we present the first comprehensive solution for dense video captioning (DVC) of surgical videos, addressing this task despite the absence of existing DVC datasets in the surgical domain. CONCLUSION: We introduce a novel approach to surgical workflow understanding that leverages video-language pretraining, large-scale video pretraining, and optimized fine-tuning. Our method improves performance over state-of-the-art techniques and enables new downstream tasks for surgical video understanding.

9A Review on the Current Applications of Artificial Intelligence in the Operating Room.PubMed

David C Birkhoff, Anne Sophie H M van Dalen, Marlies P Schijven
. Artificial intelligence (AI) is an era upcoming in medicine and, more recently, in the operating room (OR). Existing literature elaborates mainly on the future possibilities and expectations for AI in surgery. The aim of this study is to systematically provide an overview of the current actual AI applications used to support processes inside the OR. . PubMed, Embase, Cochrane Library, and IEEE Xplore were searched using inclusion criteria for relevant articles up to August 25th, 2020. No study types were excluded beforehand. Articles describing current AI applications for surgical purposes inside the OR were reviewed. . Nine studies were included. An overview of the researched and described applications of AI in the OR is provided, including procedure duration prediction, gesture recognition, intraoperative cancer detection, intraoperative video analysis, workflow recognition, an endoscopic guidance system, knot-tying, and automatic registration and tracking of the bone in orthopedic surgery. These technologies are compared to their, often non-AI, baseline alternatives. . Currently described applications of AI in the OR are limited to date. They may, however, have a promising future in improving surgical precision, reduce manpower, support intraoperative decision-making, and increase surgical safety. Nonetheless, the application and implementation of AI inside the OR still has several challenges to overcome. Clear regulatory, organizational, and clinical conditions are imperative for AI to redeem its promise. Future research on use of AI in the OR should therefore focus on clinical validation of AI applications, the legal and ethical considerations, and on evaluation of implementation trajectory.

10Artificial Intelligence, Machine Learning, and Surgical Science: Reality Versus Hype.PubMed

Majed El Hechi, Thomas M Ward, Gary C An, et al.
Artificial intelligence (AI) has made increasing inroads in clinical medicine. In surgery, machine learning-based algorithms are being studied for use as decision aids in risk prediction and even for intraoperative applications, including image recognition and video analysis. While AI has great promise in surgery, these algorithms come with a series of potential pitfalls that cannot be ignored as hospital systems and surgeons consider implementing these technologies. The aim of this review is to discuss the progress, promise, and pitfalls of AI in surgery.
内容由 AI 生成,仅供参考,请仔细甄别