• 文献检索
  • 文档翻译
  • 深度研究
  • 学术资讯
  • Suppr Zotero 插件Zotero 插件
  • 邀请有礼
  • 套餐&价格
  • 历史记录
应用&插件
Suppr Zotero 插件Zotero 插件浏览器插件Mac 客户端Windows 客户端微信小程序
定价
高级版会员购买积分包购买API积分包
服务
文献检索文档翻译深度研究API 文档MCP 服务
关于我们
关于 Suppr公司介绍联系我们用户协议隐私条款
关注我们

Suppr 超能文献

核心技术专利:CN118964589B侵权必究
粤ICP备2023148730 号-1Suppr @ 2026

文献检索

告别复杂PubMed语法,用中文像聊天一样搜索,搜遍4000万医学文献。AI智能推荐,让科研检索更轻松。

立即免费搜索

文件翻译

保留排版,准确专业,支持PDF/Word/PPT等文件格式,支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述,25分钟生成高质量综述,智能提取关键信息,辅助科研写作。

立即免费体验

用于稳健且增量式数据可视化的自组织星云状生长

Self-Organizing Nebulous Growths for Robust and Incremental Data Visualization.

作者信息

Senanayake Damith A, Wang Wei, Naik Shalin H, Halgamuge Saman

出版信息

IEEE Trans Neural Netw Learn Syst. 2021 Oct;32(10):4588-4602. doi: 10.1109/TNNLS.2020.3023941. Epub 2021 Oct 5.

DOI:10.1109/TNNLS.2020.3023941
PMID:32997635
Abstract

Nonparametric dimensionality reduction techniques, such as t-distributed Stochastic Neighbor Embedding (t-SNE) and uniform manifold approximation and projection (UMAP), are proficient in providing visualizations for data sets of fixed sizes. However, they cannot incrementally map and insert new data points into an already provided data visualization. We present self-organizing nebulous growths (SONG), a parametric nonlinear dimensionality reduction technique that supports incremental data visualization, i.e., incremental addition of new data while preserving the structure of the existing visualization. In addition, SONG is capable of handling new data increments, no matter whether they are similar or heterogeneous to the already observed data distribution. We test SONG on a variety of real and simulated data sets. The results show that SONG is superior to Parametric t-SNE, t-SNE, and UMAP in incremental data visualization. Especially, for heterogeneous increments, SONG improves over Parametric t-SNE by 14.98% on the Fashion MNIST data set and 49.73% on the MNIST data set regarding the cluster quality measured by the adjusted mutual information scores. On similar or homogeneous increments, the improvements are 8.36% and 42.26%, respectively. Furthermore, even when the abovementioned data sets are presented all at once, SONG performs better or comparable to UMAP and superior to t-SNE. We also demonstrate that the algorithmic foundations of SONG render it more tolerant to noise compared with UMAP and t-SNE, thus providing greater utility for data with high variance, high mixing of clusters, or noise.

摘要

非参数降维技术,如t分布随机邻域嵌入(t-SNE)和均匀流形近似与投影(UMAP),擅长为固定大小的数据集提供可视化。然而,它们无法将新数据点增量映射并插入到已有的数据可视化中。我们提出了自组织星云增长(SONG),这是一种参数化非线性降维技术,支持增量数据可视化,即在保留现有可视化结构的同时增量添加新数据。此外,SONG能够处理新的数据增量,无论它们与已观察到的数据分布相似还是异构。我们在各种真实和模拟数据集上测试了SONG。结果表明,在增量数据可视化方面,SONG优于参数化t-SNE、t-SNE和UMAP。特别是,对于异构增量,在Fashion MNIST数据集上,SONG相对于参数化t-SNE在通过调整互信息分数衡量的聚类质量方面提高了14.98%,在MNIST数据集上提高了49.73%。在相似或同构增量方面,改进分别为8.36%和42.26%。此外,即使上述数据集一次性全部呈现,SONG的表现也优于或与UMAP相当,且优于t-SNE。我们还证明,与UMAP和t-SNE相比,SONG的算法基础使其对噪声更具容忍性,从而为具有高方差、高聚类混合或噪声的数据提供了更大的实用性。

相似文献

1
Self-Organizing Nebulous Growths for Robust and Incremental Data Visualization.用于稳健且增量式数据可视化的自组织星云状生长
IEEE Trans Neural Netw Learn Syst. 2021 Oct;32(10):4588-4602. doi: 10.1109/TNNLS.2020.3023941. Epub 2021 Oct 5.
2
Evaluation of Distance Metrics and Spatial Autocorrelation in Uniform Manifold Approximation and Projection Applied to Mass Spectrometry Imaging Data.基于均摊近似和投影的距离度量和空间自相关评估及其在质谱成像数据中的应用。
Anal Chem. 2019 May 7;91(9):5706-5714. doi: 10.1021/acs.analchem.8b05827. Epub 2019 Apr 25.
3
Dimensionality reduction by UMAP reinforces sample heterogeneity analysis in bulk transcriptomic data.UMAP 通过降维增强了批量转录组数据中样本异质性分析。
Cell Rep. 2021 Jul 27;36(4):109442. doi: 10.1016/j.celrep.2021.109442.
4
Shape-aware stochastic neighbor embedding for robust data visualisations.形状感知随机近邻嵌入的稳健数据可视化。
BMC Bioinformatics. 2022 Nov 14;23(1):477. doi: 10.1186/s12859-022-05028-8.
5
A cross entropy test allows quantitative statistical comparison of t-SNE and UMAP representations.交叉熵测试允许对 t-SNE 和 UMAP 表示进行定量统计比较。
Cell Rep Methods. 2023 Jan 13;3(1):100390. doi: 10.1016/j.crmeth.2022.100390. eCollection 2023 Jan 23.
6
DGCyTOF: Deep learning with graphic cluster visualization to predict cell types of single cell mass cytometry data.DGCyTOF:基于图形聚类可视化的深度学习,用于预测单细胞质谱流式细胞术数据的细胞类型。
PLoS Comput Biol. 2022 Apr 11;18(4):e1008885. doi: 10.1371/journal.pcbi.1008885. eCollection 2022 Apr.
7
Visualizing Single-Cell RNA-seq Data with Semisupervised Principal Component Analysis.基于半监督主成分分析的单细胞 RNA-seq 数据可视化
Int J Mol Sci. 2020 Aug 12;21(16):5797. doi: 10.3390/ijms21165797.
8
Using Global t-SNE to Preserve Intercluster Data Structure.使用全局 t-SNE 保持簇间数据结构。
Neural Comput. 2022 Jul 14;34(8):1637-1651. doi: 10.1162/neco_a_01504.
9
Self-Organizing Map and Relational Perspective Mapping for the Accurate Visualization of High-Dimensional Hyperspectral Data.自组织映射和关系透视映射在高维高光谱数据的精确可视化中的应用。
Anal Chem. 2020 Aug 4;92(15):10450-10459. doi: 10.1021/acs.analchem.0c00986. Epub 2020 Jul 16.
10
Statistical method scDEED for detecting dubious 2D single-cell embeddings and optimizing t-SNE and UMAP hyperparameters.用于检测可疑的 2D 单细胞嵌入并优化 t-SNE 和 UMAP 参数的统计方法 scDEED。
Nat Commun. 2024 Feb 26;15(1):1753. doi: 10.1038/s41467-024-45891-y.

引用本文的文献

1
Advancing hyperspectral imaging and machine learning tools toward clinical adoption in tissue diagnostics: A comprehensive review.推动高光谱成像和机器学习工具在组织诊断中的临床应用:全面综述。
APL Bioeng. 2024 Dec 6;8(4):041504. doi: 10.1063/5.0240444. eCollection 2024 Dec.
2
Visualization of incrementally learned projection trajectories for longitudinal data.纵向数据中渐进式学习投影轨迹的可视化。
Sci Rep. 2024 Jun 12;14(1):13558. doi: 10.1038/s41598-024-63511-z.