ClusterMine：一种基于基因集表达谱的知识整合聚类方法。

ClusterMine: A knowledge-integrated clustering approach based on expression profiles of gene sets.

机构信息

Hunan Provincial Key Lab on Bioinformatics, School of Computer Science and Engineering, Central South University, Changsha 400083, P. R. China.

School of Computer Science and Engineering, Yulin Normal University, Yulin, Guangxi, P. R. China.

出版信息

J Bioinform Comput Biol. 2020 Jun;18(3):2040009. doi: 10.1142/S0219720020400090.

DOI:10.1142/S0219720020400090

PMID:32698720

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC8864677/

Abstract

Clustering analysis of gene expression data is essential for understanding complex biological data, and is widely used in important biological applications such as the identification of cell subpopulations and disease subtypes. In commonly used methods such as hierarchical clustering (HC) and consensus clustering (CC), holistic expression profiles of all genes are often used to assess the similarity between samples for clustering. While these methods have been proven successful in identifying sample clusters in many areas, they do not provide information about which gene sets (functions) contribute most to the clustering, thus limiting the interpretability of the resulting cluster. We hypothesize that integrating prior knowledge of annotated gene sets would not only achieve satisfactory clustering performance but also, more importantly, enable potential biological interpretation of clusters. Here we report ClusterMine, an approach that identifies clusters by assessing functional similarity between samples through integrating known annotated gene sets in functional annotation databases such as Gene Ontology. In addition to the cluster membership of each sample as provided by conventional approaches, it also outputs gene sets that most likely contribute to the clustering, thus facilitating biological interpretation. We compare ClusterMine with conventional approaches on nine real-world experimental datasets that represent different application scenarios in biology. We find that ClusterMine achieves better performances and that the gene sets prioritized by our method are biologically meaningful. ClusterMine is implemented as an R package and is freely available at: www.genemine.org/clustermine.php.

摘要

基因表达数据的聚类分析对于理解复杂的生物数据至关重要，广泛应用于细胞亚群和疾病亚型识别等重要的生物学应用中。在层次聚类（HC）和共识聚类（CC）等常用方法中，通常使用所有基因的整体表达谱来评估样本之间的相似性以进行聚类。虽然这些方法在许多领域成功地识别了样本聚类，但它们没有提供哪些基因集（功能）对聚类的贡献最大的信息，从而限制了聚类结果的可解释性。我们假设整合已知注释基因集的先验知识不仅可以达到令人满意的聚类性能，而且更重要的是，可以对聚类进行潜在的生物学解释。在这里，我们报告了 ClusterMine，这是一种通过整合功能注释数据库（如基因本体论）中的已知注释基因集来评估样本之间功能相似性以识别聚类的方法。除了传统方法提供的每个样本的聚类成员身份外，它还输出最有可能导致聚类的基因集，从而促进生物学解释。我们在九个代表生物学不同应用场景的真实实验数据集上比较了 ClusterMine 和传统方法。我们发现 ClusterMine 具有更好的性能，并且我们方法优先的基因集具有生物学意义。ClusterMine 作为 R 包实现，并可在：www.genemine.org/clustermine.php 免费获得。

相似文献

ClusterMine: A knowledge-integrated clustering approach based on expression profiles of gene sets.ClusterMine：一种基于基因集表达谱的知识整合聚类方法。

J Bioinform Comput Biol. 2020 Jun;18(3):2040009. doi: 10.1142/S0219720020400090.

Beyond synexpression relationships: local clustering of time-shifted and inverted gene expression profiles identifies new, biologically relevant interactions.超越共表达关系：时移和反向基因表达谱的局部聚类可识别新的生物学相关相互作用。

J Mol Biol. 2001 Dec 14;314(5):1053-66. doi: 10.1006/jmbi.2000.5219.

GO functional similarity clustering depends on similarity measure, clustering method, and annotation completeness.GO 功能相似性聚类取决于相似性度量、聚类方法和注释完整性。

BMC Bioinformatics. 2019 Mar 27;20(1):155. doi: 10.1186/s12859-019-2752-2.

SC(3): Triple spectral clustering-based consensus clustering framework for class discovery from cancer gene expression profiles.SC(3)：基于三重谱聚类的共识聚类框架，用于从癌症基因表达谱中发现类别。

IEEE/ACM Trans Comput Biol Bioinform. 2012 Nov-Dec;9(6):1751-65. doi: 10.1109/TCBB.2012.108.

Dynamically weighted clustering with noise set.带噪声集的动态加权聚类。

Bioinformatics. 2010 Feb 1;26(3):341-7. doi: 10.1093/bioinformatics/btp671. Epub 2009 Dec 9.

Adaptive Fuzzy Consensus Clustering Framework for Clustering Analysis of Cancer Data.用于癌症数据聚类分析的自适应模糊共识聚类框架

IEEE/ACM Trans Comput Biol Bioinform. 2015 Jul-Aug;12(4):887-901. doi: 10.1109/TCBB.2014.2359433.

A network-assisted co-clustering algorithm to discover cancer subtypes based on gene expression.基于基因表达的网络辅助协同聚类算法发现癌症亚型。

BMC Bioinformatics. 2014 Feb 4;15:37. doi: 10.1186/1471-2105-15-37.

Knowledge-assisted recognition of cluster boundaries in gene expression data.基因表达数据中聚类边界的知识辅助识别。

Artif Intell Med. 2005 Sep-Oct;35(1-2):171-83. doi: 10.1016/j.artmed.2005.02.007.

Fuzzy c-means clustering with prior biological knowledge.具有先验生物学知识的模糊c均值聚类

J Biomed Inform. 2009 Feb;42(1):74-81. doi: 10.1016/j.jbi.2008.05.009. Epub 2008 May 24.

Judging the quality of gene expression-based clustering methods using gene annotation.利用基因注释评估基于基因表达的聚类方法的质量。

Genome Res. 2002 Oct;12(10):1574-81. doi: 10.1101/gr.397002.

引用本文的文献

The Cytotoxic Properties of Extreme Fungi's Bioactive Components-An Updated Metabolic and Omics Overview.极端真菌生物活性成分的细胞毒性特性——代谢与组学最新综述

Life (Basel). 2023 Jul 25;13(8):1623. doi: 10.3390/life13081623.

SSRE: Cell Type Detection Based on Sparse Subspace Representation and Similarity Enhancement.SSRE：基于稀疏子空间表示和相似度增强的细胞类型检测。

Genomics Proteomics Bioinformatics. 2021 Apr;19(2):282-291. doi: 10.1016/j.gpb.2020.09.004. Epub 2021 Feb 27.

本文引用的文献

Schizophrenia Identification Using Multi-View Graph Measures of Functional Brain Networks.使用功能性脑网络的多视图图测度识别精神分裂症

Front Bioeng Biotechnol. 2020 Jan 15;7:479. doi: 10.3389/fbioe.2019.00479. eCollection 2019.

Drug repositioning based on bounded nuclear norm regularization.基于有界核范数正则化的药物重定位。

Bioinformatics. 2019 Jul 15;35(14):i455-i463. doi: 10.1093/bioinformatics/btz331.

A Gene Rank Based Approach for Single Cell Similarity Assessment and Clustering.基于基因排序的单细胞相似性评估和聚类方法。

IEEE/ACM Trans Comput Biol Bioinform. 2021 Mar-Apr;18(2):431-442. doi: 10.1109/TCBB.2019.2931582. Epub 2021 Apr 6.

DeepSignal: detecting DNA methylation state from Nanopore sequencing reads using deep-learning.DeepSignal：使用深度学习从纳米孔测序reads 中检测 DNA 甲基化状态。

Bioinformatics. 2019 Nov 1;35(22):4586-4595. doi: 10.1093/bioinformatics/btz276.

SinNLRR: a robust subspace clustering method for cell type detection by non-negative and low-rank representation.SinNLRR：一种基于非负低秩表示的稳健子空间聚类方法，用于细胞类型检测。

Bioinformatics. 2019 Oct 1;35(19):3642-3650. doi: 10.1093/bioinformatics/btz139.

Single-cell RNA-seq denoising using a deep count autoencoder.基于深度计数自编码器的单细胞 RNA-seq 去噪。

Nat Commun. 2019 Jan 23;10(1):390. doi: 10.1038/s41467-018-07931-2.

Machine learning and feature selection for drug response prediction in precision oncology applications.精准肿瘤学应用中用于药物反应预测的机器学习与特征选择

Biophys Rev. 2019 Feb;11(1):31-39. doi: 10.1007/s12551-018-0446-z. Epub 2018 Aug 10.

A comprehensive evaluation of module detection methods for gene expression data.基因表达数据模块检测方法的综合评估

Nat Commun. 2018 Mar 15;9(1):1090. doi: 10.1038/s41467-018-03424-4.

Regulation of Cell Cycle to Stimulate Adult Cardiomyocyte Proliferation and Cardiac Regeneration.调控细胞周期以刺激成体心肌细胞增殖和心脏再生。

Cell. 2018 Mar 22;173(1):104-116.e12. doi: 10.1016/j.cell.2018.02.014. Epub 2018 Mar 1.

Spectral clustering based on learning similarity matrix.基于学习相似性矩阵的谱聚类。

Bioinformatics. 2018 Jun 15;34(12):2069-2076. doi: 10.1093/bioinformatics/bty050.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验