Suppr超能文献

谱聚类、贝叶斯生成森林和森林过程。

Spectral Clustering, Bayesian Spanning Forest, and Forest Process.

作者信息

Duan Leo L, Roy Arkaprava

机构信息

Department of Statistics, University of Florida.

Department of Biostatistics, University of Florida.

出版信息

J Am Stat Assoc. 2024;119(547):2140-2153. doi: 10.1080/01621459.2023.2250098. Epub 2023 Sep 29.

Abstract

Spectral clustering views the similarity matrix as a weighted graph, and partitions the data by minimizing a graph-cut loss. Since it minimizes the across-cluster similarity, there is no need to model the distribution within each cluster. As a result, one reduces the chance of model misspecification, which is often a risk in mixture model-based clustering. Nevertheless, compared to the latter, spectral clustering has no direct ways of quantifying the clustering uncertainty (such as the assignment probability), or allowing easy model extensions for complicated data applications. To fill this gap, we propose the Bayesian forest model as a generative graphical model for spectral clustering. This is motivated by our discovery that the posterior connecting matrix in a forest model has almost the same leading eigenvectors, as the ones used by normalized spectral clustering. To induce a distribution for the forest, we develop a "forest process" as a graph extension to the urn process, while we carefully characterize the differences in the partition probability. We derive a simple Markov chain Monte Carlo algorithm for posterior estimation, and demonstrate superior performance compared to existing algorithms. We illustrate several model-based extensions useful for data applications, including high-dimensional and multi-view clustering for images.

摘要

谱聚类将相似性矩阵视为加权图,并通过最小化图割损失来对数据进行划分。由于它最小化了簇间相似性,因此无需对每个簇内的分布进行建模。这样一来,就降低了模型误设的可能性,而在基于混合模型的聚类中,模型误设往往是一个风险。然而,与后者相比,谱聚类没有直接量化聚类不确定性的方法(如分配概率),也不便于对复杂的数据应用进行模型扩展。为了填补这一空白,我们提出将贝叶斯森林模型作为谱聚类的生成式图形模型。这是基于我们的发现:森林模型中的后验连接矩阵具有几乎与归一化谱聚类所使用的相同的主特征向量。为了诱导森林的分布,我们开发了一种“森林过程”,作为对瓮过程的图形扩展,同时仔细刻画了划分概率的差异。我们推导了一种用于后验估计的简单马尔可夫链蒙特卡罗算法,并证明了其与现有算法相比具有优越的性能。我们展示了几种对数据应用有用的基于模型的扩展,包括用于图像的高维聚类和多视图聚类。

相似文献

1
Spectral Clustering, Bayesian Spanning Forest, and Forest Process.谱聚类、贝叶斯生成森林和森林过程。
J Am Stat Assoc. 2024;119(547):2140-2153. doi: 10.1080/01621459.2023.2250098. Epub 2023 Sep 29.
2
Towards a unified framework for graph-based multi-view clustering.面向基于图的多视图聚类的统一框架。
Neural Netw. 2024 May;173:106197. doi: 10.1016/j.neunet.2024.106197. Epub 2024 Feb 23.
7
Balance guided incomplete multi-view spectral clustering.平衡引导的不完全多视图谱聚类。
Neural Netw. 2023 Sep;166:260-272. doi: 10.1016/j.neunet.2023.07.022. Epub 2023 Jul 20.

本文引用的文献

1
Bias-adjusted spectral clustering in multi-layer stochastic block models.多层随机块模型中的偏差调整谱聚类
J Am Stat Assoc. 2023;118(544):2433-2445. doi: 10.1080/01621459.2022.2054817. Epub 2022 Apr 25.
3
Bayesian Distance Clustering.贝叶斯距离聚类
J Mach Learn Res. 2021 Jan-Dec;22.
6
Robust Bayesian inference via coarsening.通过粗化进行稳健贝叶斯推断。
J Am Stat Assoc. 2019;114(527):1113-1125. doi: 10.1080/01621459.2018.1469995. Epub 2018 Aug 6.
7
Automated anatomical labelling atlas 3.自动解剖学标注图谱 3.
Neuroimage. 2020 Feb 1;206:116189. doi: 10.1016/j.neuroimage.2019.116189. Epub 2019 Sep 12.
8
Mixture models with a prior on the number of components.对组件数量具有先验的混合模型。
J Am Stat Assoc. 2018;113(521):340-356. doi: 10.1080/01621459.2016.1255636. Epub 2017 Nov 13.
9
Identifying Mixtures of Mixtures Using Bayesian Estimation.使用贝叶斯估计识别混合混合物。
J Comput Graph Stat. 2017 Apr 3;26(2):285-295. doi: 10.1080/10618600.2016.1200472. Epub 2017 Apr 24.

文献AI研究员

20分钟写一篇综述,助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型,支持多种主流文档格式。

立即体验