使用开源大型语言模型对医疗保健定性访谈进行归纳主题分析：与传统方法相比如何？

Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?

机构信息

Department of Psychiatry, Yale University School of Medicine, New Haven, CT, USA.

出版信息

Comput Methods Programs Biomed. 2024 Oct;255:108356. doi: 10.1016/j.cmpb.2024.108356. Epub 2024 Jul 24.

DOI:10.1016/j.cmpb.2024.108356

PMID:39067136

Abstract

BACKGROUND

Large language models (LLMs) are generative artificial intelligence that have ignited much interest and discussion about their utility in clinical and research settings. Despite this interest there is sparse analysis of their use in qualitative thematic analysis comparing their current ability to that of human coding and analysis. In addition, there has been no published analysis of their use in real-world, protected health information.

OBJECTIVE

Here we fill that gap in the literature by comparing an LLM to standard human thematic analysis in real-world, semi-structured interviews of both patients and clinicians within a psychiatric setting.

METHODS

Using a 70 billion parameter open-source LLM running on local hardware and advanced prompt engineering techniques, we produced themes that summarized a full corpus of interviews in minutes. Subsequently we used three different evaluation methods for quantifying similarity between themes produced by the LLM and those produced by humans.

RESULTS

These revealed similarities ranging from moderate to substantial (Jaccard similarity coefficients 0.44-0.69), which are promising preliminary results.

CONCLUSION

Our study demonstrates that open-source LLMs can effectively generate robust themes from qualitative data, achieving substantial similarity to human-generated themes. The validation of LLMs in thematic analysis, coupled with evaluation methodologies, highlights their potential to enhance and democratize qualitative research across diverse fields.

摘要

背景

大型语言模型（LLMs）是生成式人工智能，它们在临床和研究环境中的实用性引起了广泛关注和讨论。尽管人们对此很感兴趣，但很少有分析将其与人类编码和分析进行比较，以评估其在定性主题分析中的应用。此外，也没有关于其在真实的、受保护的健康信息中的应用的已发表分析。

目的

本研究通过在精神科环境中对患者和临床医生进行的真实半结构化访谈，将 LLM 与标准的人类主题分析进行比较，填补了这一文献空白。

方法

我们使用一个 700 亿参数的开源 LLM，在本地硬件上运行，并采用先进的提示工程技术，在数分钟内生成了对整个访谈语料库的总结主题。随后，我们使用三种不同的评估方法来量化 LLM 生成的主题与人类生成的主题之间的相似性。

结果

这些结果显示出相似性从中等到高度（Jaccard 相似系数为 0.44-0.69），这是有希望的初步结果。

结论

我们的研究表明，开源 LLM 可以有效地从定性数据中生成强大的主题，与人类生成的主题具有高度相似性。对 LLM 在主题分析中的验证，以及评估方法的使用，突出了它们在不同领域增强和民主化定性研究的潜力。

相似文献

Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?使用开源大型语言模型对医疗保健定性访谈进行归纳主题分析：与传统方法相比如何？

Comput Methods Programs Biomed. 2024 Oct;255:108356. doi: 10.1016/j.cmpb.2024.108356. Epub 2024 Jul 24.

Large Language Models Can Enable Inductive Thematic Analysis of a Social Media Corpus in a Single Prompt: Human Validation Study.大语言模型可通过单一提示实现社交媒体语料库的归纳主题分析：人类验证研究。

JMIR Infodemiology. 2024 Aug 29;4:e59641. doi: 10.2196/59641.

Comparing the Efficacy and Efficiency of Human and Generative AI: Qualitative Thematic Analyses.比较人类与生成式人工智能的功效和效率：定性主题分析

JMIR AI. 2024 Aug 2;3:e54482. doi: 10.2196/54482.

Clinician voices on ethics of LLM integration in healthcare: a thematic analysis of ethical concerns and implications.临床医生对医疗保健中 LLM 整合的伦理看法：对伦理问题和影响的主题分析。

BMC Med Inform Decis Mak. 2024 Sep 9;24(1):250. doi: 10.1186/s12911-024-02656-3.

Applying Large Language Models to Interpret Qualitative Interviews in Healthcare.将大型语言模型应用于医疗保健中的定性访谈解释。

Stud Health Technol Inform. 2024 Aug 22;316:791-795. doi: 10.3233/SHTI240530.

Quality of Answers of Generative Large Language Models Versus Peer Users for Interpreting Laboratory Test Results for Lay Patients: Evaluation Study.生成式大语言模型与同行用户对解释非专业患者实验室检测结果的答案质量比较：评估研究。

J Med Internet Res. 2024 Apr 17;26:e56655. doi: 10.2196/56655.

The Role of Large Language Models in Transforming Emergency Medicine: Scoping Review.大型语言模型在变革急诊医学中的作用：范围综述

JMIR Med Inform. 2024 May 10;12:e53787. doi: 10.2196/53787.

Framework-based qualitative analysis of free responses of Large Language Models: Algorithmic fidelity.基于框架的大语言模型自由回答定性分析：算法保真度

PLoS One. 2024 Mar 12;19(3):e0300024. doi: 10.1371/journal.pone.0300024. eCollection 2024.

Potential of Large Language Models in Health Care: Delphi Study.大语言模型在医疗保健中的潜力：德尔菲研究。

J Med Internet Res. 2024 May 13;26:e52399. doi: 10.2196/52399.

Assessing the Alignment of Large Language Models With Human Values for Mental Health Integration: Cross-Sectional Study Using Schwartz's Theory of Basic Values.评估大型语言模型与人类心理健康整合价值观的一致性：使用施瓦茨基本价值观理论的横断面研究。

JMIR Ment Health. 2024 Apr 9;11:e55988. doi: 10.2196/55988.

引用本文的文献

Frankenstein, thematic analysis and generative artificial intelligence: Quality appraisal methods and considerations for qualitative research.《科学怪人》、主题分析与生成式人工智能：定性研究的质量评估方法及考量

PLoS One. 2025 Sep 5;20(9):e0330217. doi: 10.1371/journal.pone.0330217. eCollection 2025.

Leveraging large language models for automated depression screening.利用大语言模型进行自动抑郁症筛查。

PLOS Digit Health. 2025 Jul 28;4(7):e0000943. doi: 10.1371/journal.pdig.0000943. eCollection 2025 Jul.

Lunch box and fruits as a simulator for teaching basic physics of ultrasound: A mixed research methods study.午餐盒与水果作为超声基础物理教学模拟器：一项混合研究方法的研究

World J Emerg Surg. 2025 Jul 25;20(1):64. doi: 10.1186/s13017-025-00637-z.

Utilizing AI-Powered Thematic Analysis: Methodology, Implementation, and Lessons Learned.利用人工智能驱动的主题分析：方法、实施与经验教训。

Cureus. 2025 Jun 4;17(6):e85338. doi: 10.7759/cureus.85338. eCollection 2025 Jun.

Effect of a generative artificial intelligence digital scribe on pediatric provider documentation time, cognitive burden, and burnout.生成式人工智能数字书记员对儿科医疗服务提供者的文档记录时间、认知负担和职业倦怠的影响。

JAMIA Open. 2025 Jul 3;8(4):ooaf068. doi: 10.1093/jamiaopen/ooaf068. eCollection 2025 Aug.

Responsive population-based cohorts as platforms for characterising pathogen- and population-level infection dynamics for epidemic prevention, preparedness and response.基于人群的反应性队列作为表征病原体和人群水平感染动态以进行疫情预防、防范和应对的平台。

Euro Surveill. 2025 Jun;30(25). doi: 10.2807/1560-7917.ES.2025.30.25.2400255.

Annotation of biological samples data to standard ontologies with support from large language models.在大语言模型的支持下将生物样本数据注释到标准本体中。

Comput Struct Biotechnol J. 2025 May 26;27:2155-2167. doi: 10.1016/j.csbj.2025.05.020. eCollection 2025.

Generative AI for thematic analysis in a maternal health study: coding semistructured interviews using large language models.生成式人工智能在孕产妇健康研究中的主题分析：使用大语言模型对半结构化访谈进行编码

Appl Psychol Health Well Being. 2025 Jun;17(3):e70038. doi: 10.1111/aphw.70038.

Generative AI for Thematic Analysis in a Maternal Health Study: Coding Semi-structured Interviews using Large Language Models (LLMs).生成式人工智能在孕产妇健康研究中的主题分析：使用大语言模型对半结构化访谈进行编码

medRxiv. 2025 Apr 23:2024.09.16.24313707. doi: 10.1101/2024.09.16.24313707.

Delving into the Practical Applications and Pitfalls of Large Language Models in Medical Education: Narrative Review.深入探讨大语言模型在医学教育中的实际应用与陷阱：叙述性综述

Adv Med Educ Pract. 2025 Apr 18;16:625-636. doi: 10.2147/AMEP.S497020. eCollection 2025.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

使用开源大型语言模型对医疗保健定性访谈进行归纳主题分析：与传统方法相比如何？

Inductive thematic analysis of healthcare qualitative interviews using open-source large language models: How does it compare to traditional methods?

机构信息

出版信息

BACKGROUND

OBJECTIVE

METHODS

RESULTS

CONCLUSION

背景

目的

方法

结果

结论

相似文献

引用本文的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

引用本文的文献