MAQC-II 乳腺癌和多发性骨髓瘤基因表达数据的特征选择和分类。

Feature selection and classification of MAQC-II breast cancer and multiple myeloma microarray gene expression data.

机构信息

Department of Computer Science and Institute for Complex Additive Systems Analysis, New Mexico Tech, Socorro, New Mexico, United States of America.

出版信息

PLoS One. 2009 Dec 11;4(12):e8250. doi: 10.1371/journal.pone.0008250.

DOI:10.1371/journal.pone.0008250

PMID:20011240

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC2789385/

Abstract

Microarray data has a high dimension of variables but available datasets usually have only a small number of samples, thereby making the study of such datasets interesting and challenging. In the task of analyzing microarray data for the purpose of, e.g., predicting gene-disease association, feature selection is very important because it provides a way to handle the high dimensionality by exploiting information redundancy induced by associations among genetic markers. Judicious feature selection in microarray data analysis can result in significant reduction of cost while maintaining or improving the classification or prediction accuracy of learning machines that are employed to sort out the datasets. In this paper, we propose a gene selection method called Recursive Feature Addition (RFA), which combines supervised learning and statistical similarity measures. We compare our method with the following gene selection methods: Support Vector Machine Recursive Feature Elimination (SVMRFE), Leave-One-Out Calculation Sequential Forward Selection (LOOCSFS), Gradient based Leave-one-out Gene Selection (GLGS). To evaluate the performance of these gene selection methods, we employ several popular learning classifiers on the MicroArray Quality Control phase II on predictive modeling (MAQC-II) breast cancer dataset and the MAQC-II multiple myeloma dataset. Experimental results show that gene selection is strictly paired with learning classifier. Overall, our approach outperforms other compared methods. The biological functional analysis based on the MAQC-II breast cancer dataset convinced us to apply our method for phenotype prediction. Additionally, learning classifiers also play important roles in the classification of microarray data and our experimental results indicate that the Nearest Mean Scale Classifier (NMSC) is a good choice due to its prediction reliability and its stability across the three performance measurements: Testing accuracy, MCC values, and AUC errors.

摘要

微阵列数据具有很高的变量维度，但可用的数据集通常只有少量的样本，因此研究此类数据集非常有趣且具有挑战性。在分析微阵列数据的任务中，例如预测基因-疾病关联，特征选择非常重要，因为它提供了一种利用遗传标记之间的关联所产生的信息冗余来处理高维数据的方法。在微阵列数据分析中进行明智的特征选择可以显著降低成本，同时保持或提高用于整理数据集的学习机器的分类或预测准确性。在本文中，我们提出了一种称为递归特征添加（RFA）的基因选择方法，该方法结合了监督学习和统计相似性度量。我们将我们的方法与以下基因选择方法进行了比较：支持向量机递归特征消除（SVMRFE）、留一计算顺序前向选择（LOOCSFS）、基于梯度的留一基因选择（GLGS）。为了评估这些基因选择方法的性能，我们在 MicroArray Quality Control phase II on predictive modeling（MAQC-II）乳腺癌数据集和 MAQC-II 多发性骨髓瘤数据集上使用了几种流行的学习分类器。实验结果表明，基因选择与学习分类器严格配对。总的来说，我们的方法优于其他比较方法。基于 MAQC-II 乳腺癌数据集的生物学功能分析使我们相信可以应用我们的方法进行表型预测。此外，学习分类器在微阵列数据的分类中也起着重要作用，我们的实验结果表明，由于其预测可靠性及其在三个性能测量中的稳定性，最近均值尺度分类器（NMSC）是一个不错的选择：测试准确性、MCC 值和 AUC 误差。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/06b4/2789385/7ae133ce510d/pone.0008250.g001.jpg

相似文献

Feature selection and classification of MAQC-II breast cancer and multiple myeloma microarray gene expression data.MAQC-II 乳腺癌和多发性骨髓瘤基因表达数据的特征选择和分类。

PLoS One. 2009 Dec 11;4(12):e8250. doi: 10.1371/journal.pone.0008250.

Gene selection and classification for cancer microarray data based on machine learning and similarity measures.基于机器学习和相似性度量的癌症基因芯片数据选择与分类。

BMC Genomics. 2011 Dec 23;12 Suppl 5(Suppl 5):S1. doi: 10.1186/1471-2164-12-S5-S1.

Comparison of feature selection and classification for MALDI-MS data.基质辅助激光解吸电离飞行时间质谱（MALDI-MS）数据的特征选择与分类比较

BMC Genomics. 2009 Jul 7;10 Suppl 1(Suppl 1):S3. doi: 10.1186/1471-2164-10-S1-S3.

Recursive cluster elimination (RCE) for classification and feature selection from gene expression data.用于从基因表达数据中进行分类和特征选择的递归聚类消除法（RCE）

BMC Bioinformatics. 2007 May 2;8:144. doi: 10.1186/1471-2105-8-144.

Enhancing the prediction of IDC breast cancer staging from gene expression profiles using hybrid feature selection methods and deep learning architecture.使用混合特征选择方法和深度学习架构增强从基因表达谱预测浸润性导管癌乳腺癌分期的能力。

Med Biol Eng Comput. 2023 Nov;61(11):2895-2919. doi: 10.1007/s11517-023-02892-1. Epub 2023 Aug 2.

Recursive feature selection with significant variables of support vectors.递归特征选择与支持向量的显著变量。

Comput Math Methods Med. 2012;2012:712542. doi: 10.1155/2012/712542. Epub 2012 Aug 15.

The feature selection bias problem in relation to high-dimensional gene data.与高维基因数据相关的特征选择偏差问题。

Artif Intell Med. 2016 Jan;66:63-71. doi: 10.1016/j.artmed.2015.11.001. Epub 2015 Nov 14.

Integration of RNA-Seq data with heterogeneous microarray data for breast cancer profiling.整合RNA测序数据与异质性微阵列数据用于乳腺癌分析。

BMC Bioinformatics. 2017 Nov 21;18(1):506. doi: 10.1186/s12859-017-1925-0.

Mixture classification model based on clinical markers for breast cancer prognosis.基于临床标志物的乳腺癌预后混合分类模型。

Artif Intell Med. 2010 Feb-Mar;48(2-3):129-37. doi: 10.1016/j.artmed.2009.07.008. Epub 2009 Dec 14.

Prediction potential of candidate biomarker sets identified and validated on gene expression data from multiple datasets.在来自多个数据集的基因表达数据上鉴定和验证的候选生物标志物集的预测潜力。

BMC Bioinformatics. 2007 Oct 26;8:415. doi: 10.1186/1471-2105-8-415.

引用本文的文献

Beta Distribution-Based Cross-Entropy for Feature Selection.基于贝塔分布的交叉熵用于特征选择。

Entropy (Basel). 2019 Aug 7;21(8):769. doi: 10.3390/e21080769.

A consensus multi-view multi-objective gene selection approach for improved sample classification.一种共识多视角多目标基因选择方法，用于提高样本分类。

BMC Bioinformatics. 2020 Sep 17;21(Suppl 13):386. doi: 10.1186/s12859-020-03681-5.

Building generalized linear models with ultrahigh dimensional features: A sequentially conditional approach.超高维特征的广义线性模型构建：一种序贯条件方法。

Biometrics. 2020 Mar;76(1):47-60. doi: 10.1111/biom.13122. Epub 2019 Nov 6.

Unsupervised gene selection using biological knowledge : application in sample clustering.利用生物学知识进行无监督基因选择：在样本聚类中的应用

BMC Bioinformatics. 2017 Nov 22;18(1):513. doi: 10.1186/s12859-017-1933-0.

Identifying Significant Features in Cancer Methylation Data Using Gene Pathway Segmentation.利用基因通路分割识别癌症甲基化数据中的显著特征

Cancer Inform. 2016 Sep 20;15:189-98. doi: 10.4137/CIN.S39859. eCollection 2016.

Discovering Pair-wise Synergies in Microarray Data.在微阵列数据中发现成对协同效应。

Sci Rep. 2016 Jul 29;6:30672. doi: 10.1038/srep30672.

A Review of Feature Selection and Feature Extraction Methods Applied on Microarray Data.应用于微阵列数据的特征选择与特征提取方法综述

Adv Bioinformatics. 2015;2015:198363. doi: 10.1155/2015/198363. Epub 2015 Jun 11.

Comprehensive evaluation of composite gene features in cancer outcome prediction.癌症预后预测中复合基因特征的综合评估。

Cancer Inform. 2015 Feb 24;13(Suppl 3):93-104. doi: 10.4137/CIN.S14028. eCollection 2014.

A novel strategy for gene selection of microarray data based on gene-to-class sensitivity information.一种基于基因对类别敏感性信息的微阵列数据基因选择新策略。

PLoS One. 2014 May 20;9(5):e97530. doi: 10.1371/journal.pone.0097530. eCollection 2014.

An algorithm for finding biologically significant features in microarray data based on a priori manifold learning.一种基于先验流形学习在微阵列数据中寻找生物学显著特征的算法。

PLoS One. 2014 Mar 3;9(3):e90562. doi: 10.1371/journal.pone.0090562. eCollection 2014.

本文引用的文献

The MicroArray Quality Control (MAQC)-II study of common practices for the development and validation of microarray-based predictive models.《基因芯片质量控制（MAQC）-II 研究：基于基因芯片的预测模型的开发和验证的常见实践》。

Nat Biotechnol. 2010 Aug;28(8):827-38. doi: 10.1038/nbt.1665. Epub 2010 Jul 30.

Temporal gene expression classification with regularised neural network.基于正则化神经网络的时间基因表达分类

Int J Bioinform Res Appl. 2005;1(4):399-413. doi: 10.1504/IJBRA.2005.008443.

A distribution free summarization method for Affymetrix GeneChip arrays.一种用于Affymetrix基因芯片阵列的无分布汇总方法。

Bioinformatics. 2007 Feb 1;23(3):321-7. doi: 10.1093/bioinformatics/btl609. Epub 2006 Dec 5.

Standards for systems biology.系统生物学标准。

Nat Rev Genet. 2006 Aug;7(8):593-605. doi: 10.1038/nrg1922.

Clustering microarray gene expression data using weighted Chinese restaurant process.使用加权中国餐馆过程对微阵列基因表达数据进行聚类

Bioinformatics. 2006 Aug 15;22(16):1988-97. doi: 10.1093/bioinformatics/btl284. Epub 2006 Jun 9.

Gene selection algorithms for microarray data based on least squares support vector machine.基于最小二乘支持向量机的微阵列数据基因选择算法

BMC Bioinformatics. 2006 Feb 27;7:95. doi: 10.1186/1471-2105-7-95.

Gene selection and classification of microarray data using random forest.使用随机森林进行微阵列数据的基因选择与分类

BMC Bioinformatics. 2006 Jan 6;7:3. doi: 10.1186/1471-2105-7-3.

A new algorithm for comparing and visualizing relationships between hierarchical and flat gene expression data clusterings.一种用于比较和可视化层次化与平面化基因表达数据聚类之间关系的新算法。

Bioinformatics. 2005 Nov 1;21(21):3993-9. doi: 10.1093/bioinformatics/bti644. Epub 2005 Sep 1.

Finding groups in gene expression data.在基因表达数据中寻找群组。

J Biomed Biotechnol. 2005 Jun 30;2005(2):215-25. doi: 10.1155/JBB.2005.215.

Bayesian neural network approaches to ovarian cancer identification from high-resolution mass spectrometry data.基于贝叶斯神经网络的从高分辨率质谱数据中识别卵巢癌的方法。

Bioinformatics. 2005 Jun;21 Suppl 1:i487-94. doi: 10.1093/bioinformatics/bti1030.

文献AI研究员

20分钟写一篇综述，助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型，支持多种主流文档格式。

立即体验

MAQC-II 乳腺癌和多发性骨髓瘤基因表达数据的特征选择和分类。

Feature selection and classification of MAQC-II breast cancer and multiple myeloma microarray gene expression data.

机构信息

出版信息

相似文献

引用本文的文献

本文引用的文献

文献AI研究员

用中文搜PubMed

文档翻译

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献