Revisiting Dimensionality Reduction Techniques for Visual Cluster Analysis: An Empirical Study

被引:36
|
作者
Xia, Jiazhi [1 ]
Zhang, Yuchen [1 ]
Song, Jie [1 ]
Chen, Yang [2 ]
Wang, Yunhai [3 ]
Liu, Shixia [4 ]
机构
[1] Cent South Univ, Sch Comp Sci & Engn, Changsha, Peoples R China
[2] 14 Data, Shanghai, Peoples R China
[3] Shandong Univ, Sch Comp Sci & Technol, Jinan, Peoples R China
[4] Tsinghua Univ, Sch Software, Beijing, Peoples R China
基金
中国国家自然科学基金;
关键词
Visualization; Task analysis; Principal component analysis; Measurement; Manifolds; Linearity; Visual perception; Dimensionality reduction; visual cluster analysis; perception-based evaluation; T-SNE; PROJECTION; QUALITY;
D O I
10.1109/TVCG.2021.3114694
中图分类号
TP31 [计算机软件];
学科分类号
081202 ; 0835 ;
摘要
Dimensionality Reduction (DR) techniques can generate 2D projections and enable visual exploration of cluster structures of high-dimensional datasets. However, different DR techniques would yield various patterns, which significantly affect the performance of visual cluster analysis tasks. We present the results of a user study that investigates the influence of different DR techniques on visual cluster analysis. Our study focuses on the most concerned property types, namely the linearity and locality, and evaluates twelve representative DR techniques that cover the concerned properties. Four controlled experiments were conducted to evaluate how the DR techniques facilitate the tasks of 1) cluster identification, 2) membership identification, 3) distance comparison, and 4) density comparison, respectively. We also evaluated users' subjective preference of the DR techniques regarding the quality of projected clusters. The results show that: 1) Non-linear and Local techniques are preferred in cluster identification and membership identification; 2) Linear techniques perform better than non-linear techniques in density comparison; 3) UMAP (Uniform Manifold Approximation and Projection) and t-SNE (t-Distributed Stochastic Neighbor Embedding) perform the best in cluster identification and membership identification; 4) NMF (Nonnegative Matrix Factorization) has competitive performance in distance comparison; 5) t-SNLE (t-Distributed Stochastic Neighbor Linear Embedding) has competitive performance in density comparison.
引用
收藏
页码:529 / 539
页数:11
相关论文
共 50 条
  • [21] Visual cluster separation using high-dimensional sharpened dimensionality reduction
    Kim, Youngjoo
    Telea, Alexandru C.
    Trager, Scott C.
    Roerdink, Jos B. T. M.
    INFORMATION VISUALIZATION, 2022, 21 (03) : 246 - 269
  • [22] A Comparison of Dimensionality Reduction Techniques in Virtual Screening
    Pasupa, Kitsuchart
    ARTIFICIAL INTELLIGENCE AND SOFT COMPUTING, PT II, 2013, 7895 : 297 - 308
  • [23] A Review on Dimensionality Reduction Techniques
    Huang, Xuan
    Wu, Lei
    Ye, Yinsong
    INTERNATIONAL JOURNAL OF PATTERN RECOGNITION AND ARTIFICIAL INTELLIGENCE, 2019, 33 (10)
  • [24] Comparing Dimensionality Reduction Techniques
    Nick, William
    Shelton, Joseph
    Bullock, Gina
    Esterline, Albert
    Asamene, Kassahun
    IEEE SOUTHEASTCON 2015, 2015,
  • [25] An Empirical Study of Linear Dimensionality Reduction for Judicial Predictive Models
    Liu, Zhenyu
    Chen, Huanhuan
    2018 8TH INTERNATIONAL CONFERENCE ON INFORMATION SCIENCE AND TECHNOLOGY (ICIST 2018), 2018, : 335 - 340
  • [26] Comparative analysis of dimensionality reduction techniques for cybersecurity in the SWaT dataset
    Mehmet Bozdal
    Kadir Ileri
    Ali Ozkahraman
    The Journal of Supercomputing, 2024, 80 : 1059 - 1079
  • [27] Analysis of Electricity Consumption Profiles by Means of Dimensionality Reduction Techniques
    Moran, Antonio
    Fuertes, Juan J.
    Prada, Miguel A.
    Alonso, Serafin
    Barrientos, Pablo
    Diaz, Ignacio
    ENGINEERING APPLICATIONS OF NEURAL NETWORKS, 2012, 311 : 152 - +
  • [28] Comparative analysis of dimensionality reduction techniques for cybersecurity in the SWaT dataset
    Bozdal, Mehmet
    Ileri, Kadir
    Ozkahraman, Ali
    JOURNAL OF SUPERCOMPUTING, 2024, 80 (01) : 1059 - 1079
  • [29] The Analysis of Dimensionality Reduction Techniques in Cryptographic Object Code Classification
    Wright, Jason L.
    Manic, Milos
    3RD INTERNATIONAL CONFERENCE ON HUMAN SYSTEM INTERACTION, 2010, : 157 - 162
  • [30] A Review Paper on Dimensionality Reduction Techniques
    Mulla, Faizan Riyaz
    Gupta, Anil Kumar
    JOURNAL OF PHARMACEUTICAL NEGATIVE RESULTS, 2022, 13 : 1263 - 1272