使用TCGA数据集进行乳腺癌分期的综合生物信息学和机器学习分析。

Comprehensive bioinformatics and machine learning analyses for breast cancer staging using TCGA dataset.

作者信息

Das Saurav Chandra, Tasnim Wahia, Rana Humayan Kabir, Acharjee Uzzal Kumar, Islam Md Manowarul, Khatun Rabea

机构信息

Department of Computer Science and Engineering, Jagannath University, Dhaka-1100, Bangladesh.

Department of Internet of Things and Robotics Engineering, Bangabandhu Sheikh Mujibur Rahman Digital University, Bangladesh, Kaliakair, Gazipur-1750, Bangladesh.

出版信息

Brief Bioinform. 2024 Nov 22;26(1). doi: 10.1093/bib/bbae628.

DOI:10.1093/bib/bbae628

PMID:39656775

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC11630003/

Abstract

Breast cancer is an alarming global health concern, including a vast and varied set of illnesses with different molecular characteristics. The fusion of sophisticated computational methodologies with extensive biological datasets has emerged as an effective strategy for unravelling complex patterns in cancer oncology. This research delves into breast cancer staging, classification, and diagnosis by leveraging the comprehensive dataset provided by the The Cancer Genome Atlas (TCGA). By integrating advanced machine learning algorithms with bioinformatics analysis, it introduces a cutting-edge methodology for identifying complex molecular signatures associated with different subtypes and stages of breast cancer. This study utilizes TCGA gene expression data to detect and categorize breast cancer through the application of machine learning and systems biology techniques. Researchers identified differentially expressed genes in breast cancer and analyzed them using signaling pathways, protein-protein interactions, and regulatory networks to uncover potential therapeutic targets. The study also highlights the roles of specific proteins (MYH2, MYL1, MYL2, MYH7) and microRNAs (such as hsa-let-7d-5p) that are the potential biomarkers in cancer progression founded on several analyses. In terms of diagnostic accuracy for cancer staging, the random forest method achieved 97.19%, while the XGBoost algorithm attained 95.23%. Bioinformatics and machine learning meet in this study to find potential biomarkers that influence the progression of breast cancer. The combination of sophisticated analytical methods and extensive genomic datasets presents a promising path for expanding our understanding and enhancing clinical outcomes in identifying and categorizing this intricate illness.

摘要

乳腺癌是一个令人担忧的全球健康问题，它包含一系列具有不同分子特征的广泛多样的疾病。将复杂的计算方法与大量生物数据集相结合，已成为揭示癌症肿瘤学复杂模式的有效策略。本研究利用癌症基因组图谱（TCGA）提供的综合数据集，深入探讨乳腺癌的分期、分类和诊断。通过将先进的机器学习算法与生物信息学分析相结合，它引入了一种前沿方法，用于识别与乳腺癌不同亚型和阶段相关的复杂分子特征。本研究利用TCGA基因表达数据，通过应用机器学习和系统生物学技术来检测和分类乳腺癌。研究人员在乳腺癌中鉴定出差异表达基因，并使用信号通路、蛋白质-蛋白质相互作用和调控网络对其进行分析，以发现潜在的治疗靶点。该研究还强调了特定蛋白质（MYH2、MYL1、MYL2、MYH7）和微小RNA（如hsa-let-7d-5p）的作用，这些基于多项分析是癌症进展中的潜在生物标志物。在癌症分期的诊断准确性方面，随机森林方法达到了97.19%，而XGBoost算法达到了95.23%。生物信息学和机器学习在本研究中相遇，以寻找影响乳腺癌进展的潜在生物标志物。复杂分析方法与广泛基因组数据集的结合，为扩大我们对这种复杂疾病的认识以及改善其识别和分类的临床结果提供了一条充满希望的途径。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/8426/11630003/ca2a9b0906ad/bbae628f1.jpg

相似文献

Comprehensive bioinformatics and machine learning analyses for breast cancer staging using TCGA dataset.使用TCGA数据集进行乳腺癌分期的综合生物信息学和机器学习分析。

Brief Bioinform. 2024 Nov 22;26(1). doi: 10.1093/bib/bbae628.

Tree-based machine learning algorithms identified minimal set of miRNA biomarkers for breast cancer diagnosis and molecular subtyping.基于树的机器学习算法确定了用于乳腺癌诊断和分子分型的最小 miRNA 生物标志物集。

Gene. 2018 Nov 30;677:111-118. doi: 10.1016/j.gene.2018.07.057. Epub 2018 Jul 25.

The role and machine learning analysis of mitochondrial autophagy-related gene expression in lung adenocarcinoma.线粒体自噬相关基因表达在肺腺癌中的作用及机器学习分析

Front Immunol. 2025 Apr 17;16:1509315. doi: 10.3389/fimmu.2025.1509315. eCollection 2025.

Analysis of diagnostic genes and molecular mechanisms of Crohn's disease and colon cancer based on machine learning algorithms.基于机器学习算法的克罗恩病和结肠癌诊断基因及分子机制分析

Sci Rep. 2024 Dec 30;14(1):31736. doi: 10.1038/s41598-024-82319-5.

Integrating bioinformatics and machine learning to identify AhR-related gene signatures for prognosis and tumor microenvironment modulation in melanoma.整合生物信息学和机器学习以识别与芳烃受体相关的基因特征，用于黑色素瘤的预后评估和肿瘤微环境调控。

Front Immunol. 2025 Jan 6;15:1519345. doi: 10.3389/fimmu.2024.1519345. eCollection 2024.

Integrated bioinformatics combined with machine learning to analyze shared biomarkers and pathways in psoriasis and cervical squamous cell carcinoma.综合生物信息学结合机器学习分析银屑病和宫颈鳞状细胞癌中的共享生物标志物和通路。

Front Immunol. 2024 May 28;15:1351908. doi: 10.3389/fimmu.2024.1351908. eCollection 2024.

Identification of aberrantly methylated differentially expressed genes in breast cancer by integrated bioinformatics analysis.整合生物信息学分析鉴定乳腺癌中异常甲基化差异表达基因。

J Cell Biochem. 2019 Sep;120(9):16229-16243. doi: 10.1002/jcb.28904. Epub 2019 May 12.

Machine learning and bioinformatics analysis of diagnostic biomarkers associated with the occurrence and development of lung adenocarcinoma.机器学习和生物信息学分析与肺腺癌发生发展相关的诊断生物标志物。

PeerJ. 2024 Jul 23;12:e17746. doi: 10.7717/peerj.17746. eCollection 2024.

Identifying miRNA as biomarker for breast cancer subtyping using association rule.使用关联规则识别 miRNA 作为乳腺癌亚型的生物标志物。

Comput Biol Med. 2024 Aug;178:108696. doi: 10.1016/j.compbiomed.2024.108696. Epub 2024 Jun 3.

Identification and Validation of Four Serum Biomarkers With Optimal Diagnostic and Prognostic Potential for Gastric Cancer Based on Machine Learning Algorithms.基于机器学习算法的四种具有最佳胃癌诊断和预后潜力的血清生物标志物的鉴定与验证

Cancer Med. 2025 Mar;14(6):e70659. doi: 10.1002/cam4.70659.

引用本文的文献

Machine learning and single-cell analysis uncover distinctive characteristics of CD300LG within the TNBC immune microenvironment: experimental validation.机器学习与单细胞分析揭示三阴性乳腺癌免疫微环境中CD300LG的独特特征：实验验证

Clin Exp Med. 2025 May 17;25(1):167. doi: 10.1007/s10238-025-01690-3.

本文引用的文献

Biological insights and novel biomarker discovery through deep learning approaches in breast cancer histopathology.通过深度学习方法在乳腺癌组织病理学中获得生物学见解和发现新型生物标志物

NPJ Breast Cancer. 2023 Apr 6;9(1):21. doi: 10.1038/s41523-023-00518-1.

Machine Learning Methods for Cancer Classification Using Gene Expression Data: A Review.使用基因表达数据进行癌症分类的机器学习方法：综述

Bioengineering (Basel). 2023 Jan 28;10(2):173. doi: 10.3390/bioengineering10020173.

Identification of Comorbidities, Genomic Associations, and Molecular Mechanisms for COVID-19 Using Bioinformatics Approaches.利用生物信息学方法鉴定 COVID-19 的合并症、基因组关联和分子机制。

Biomed Res Int. 2023 Jan 11;2023:6996307. doi: 10.1155/2023/6996307. eCollection 2023.

Deep Learning Based Methods for Breast Cancer Diagnosis: A Systematic Review and Future Direction.基于深度学习的乳腺癌诊断方法：系统综述与未来方向

Diagnostics (Basel). 2023 Jan 3;13(1):161. doi: 10.3390/diagnostics13010161.

Bioinformatics and System Biological Approaches for the Identification of Genetic Risk Factors in the Progression of Cardiovascular Disease.生物信息学和系统生物学方法在心血管疾病进展中遗传风险因素的鉴定。

Cardiovasc Ther. 2022 Aug 9;2022:9034996. doi: 10.1155/2022/9034996. eCollection 2022.

Determination of a six-gene prognostic model for cervical cancer based on WGCNA combined with LASSO and Cox-PH analysis.基于 WGCNA 联合 LASSO 和 Cox-PH 分析的宫颈癌六基因预后模型的建立。

World J Surg Oncol. 2021 Sep 16;19(1):277. doi: 10.1186/s12957-021-02384-2.

Transcriptome profiling by combined machine learning and statistical R analysis identifies TMEM236 as a potential novel diagnostic biomarker for colorectal cancer.联合机器学习和统计 R 分析的转录组谱分析鉴定 TMEM236 为结直肠癌的潜在新型诊断生物标志物。

Sci Rep. 2021 Jul 12;11(1):14304. doi: 10.1038/s41598-021-92692-0.

Human serum mid-infrared spectroscopy combined with machine learning algorithms for rapid detection of gliomas.人血清中红外光谱结合机器学习算法快速检测脑胶质瘤。

Photodiagnosis Photodyn Ther. 2021 Sep;35:102308. doi: 10.1016/j.pdpdt.2021.102308. Epub 2021 Apr 24.

Gene Set Knowledge Discovery with Enrichr.基因集知识发现与 Enrichr

Curr Protoc. 2021 Mar;1(3):e90. doi: 10.1002/cpz1.90.

Cancer Statistics, 2021.癌症统计数据，2021.

CA Cancer J Clin. 2021 Jan;71(1):7-33. doi: 10.3322/caac.21654. Epub 2021 Jan 12.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验

使用TCGA数据集进行乳腺癌分期的综合生物信息学和机器学习分析。

Comprehensive bioinformatics and machine learning analyses for breast cancer staging using TCGA dataset.

作者信息

机构信息

出版信息

相似文献

引用本文的文献

本文引用的文献

文献检索

文件翻译

深度研究

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献