Revisiting Dimensionality Reduction Techniques for Visual Cluster Analysis: An Empirical Study

被引:36
|
作者
Xia, Jiazhi [1 ]
Zhang, Yuchen [1 ]
Song, Jie [1 ]
Chen, Yang [2 ]
Wang, Yunhai [3 ]
Liu, Shixia [4 ]
机构
[1] Cent South Univ, Sch Comp Sci & Engn, Changsha, Peoples R China
[2] 14 Data, Shanghai, Peoples R China
[3] Shandong Univ, Sch Comp Sci & Technol, Jinan, Peoples R China
[4] Tsinghua Univ, Sch Software, Beijing, Peoples R China
基金
中国国家自然科学基金;
关键词
Visualization; Task analysis; Principal component analysis; Measurement; Manifolds; Linearity; Visual perception; Dimensionality reduction; visual cluster analysis; perception-based evaluation; T-SNE; PROJECTION; QUALITY;
D O I
10.1109/TVCG.2021.3114694
中图分类号
TP31 [计算机软件];
学科分类号
081202 ; 0835 ;
摘要
Dimensionality Reduction (DR) techniques can generate 2D projections and enable visual exploration of cluster structures of high-dimensional datasets. However, different DR techniques would yield various patterns, which significantly affect the performance of visual cluster analysis tasks. We present the results of a user study that investigates the influence of different DR techniques on visual cluster analysis. Our study focuses on the most concerned property types, namely the linearity and locality, and evaluates twelve representative DR techniques that cover the concerned properties. Four controlled experiments were conducted to evaluate how the DR techniques facilitate the tasks of 1) cluster identification, 2) membership identification, 3) distance comparison, and 4) density comparison, respectively. We also evaluated users' subjective preference of the DR techniques regarding the quality of projected clusters. The results show that: 1) Non-linear and Local techniques are preferred in cluster identification and membership identification; 2) Linear techniques perform better than non-linear techniques in density comparison; 3) UMAP (Uniform Manifold Approximation and Projection) and t-SNE (t-Distributed Stochastic Neighbor Embedding) perform the best in cluster identification and membership identification; 4) NMF (Nonnegative Matrix Factorization) has competitive performance in distance comparison; 5) t-SNLE (t-Distributed Stochastic Neighbor Linear Embedding) has competitive performance in density comparison.
引用
收藏
页码:529 / 539
页数:11
相关论文
共 50 条
  • [31] Comparing dimensionality reduction techniques for visual analysis of the LSTM hidden activity on multi-dimensional time series modeling
    Ji, Lianen
    Qiu, Shirong
    Xu, Zhi
    Liu, Yue
    Yang, Guang
    VISUAL COMPUTER, 2024, 40 (11) : 8243 - 8261
  • [32] Dimensionality reduction for the analysis of brain oscillations
    Haufe, Stefan
    Daehne, Sven
    Nikulin, Vadim V.
    NEUROIMAGE, 2014, 101 : 583 - 597
  • [33] Visual analysis of a cold rolling process using a dimensionality reduction approach
    Perez, Daniel
    Garcia-Fernandez, Francisco J.
    Diaz, Ignacio
    Cuadrado, Abel A.
    Ordonez, Daniel G.
    Diez, Alberto B.
    Dominguez, Manuel
    ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE, 2013, 26 (08) : 1865 - 1871
  • [34] A Visual Interaction Framework for Dimensionality Reduction Based Data Exploration
    Cavallo, Marco
    Demiralp, Cagatay
    PROCEEDINGS OF THE 2018 CHI CONFERENCE ON HUMAN FACTORS IN COMPUTING SYSTEMS (CHI 2018), 2018,
  • [35] Evaluating Dimensionality Reduction Techniques in Bitcoin Ransomware Detection: Comparative Analysis of Incremental PCA and UMAP
    Amissah, Daniel Kwame
    Yaokumah, Winfred
    Ansong, Edward Danso
    Appati, Justice Kwame
    SECURITY AND PRIVACY, 2025, 8 (02):
  • [36] Analysis of electricity consumption profiles in public buildings with dimensionality reduction techniques
    Moran, Antonio
    Fuertes, Juan J.
    Prada, Miguel A.
    Alonso, Serafin
    Barrientos, Pablo
    Diaz, Ignacio
    Dominguez, Manuel
    ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE, 2013, 26 (08) : 1872 - 1880
  • [37] A comparative user study of visualization techniques for cluster analysis of multidimensional data sets
    Ventocilla, Elio
    Riveiro, Maria
    INFORMATION VISUALIZATION, 2020, 19 (04) : 318 - 338
  • [38] Study of Dimensionality Reduction Techniques for Effective Investment Portfolio Data Management
    Gadre-Patwardhan, Swapnaja
    Katdare, Vivek
    Joshi, Manish
    SMART COMPUTING AND INFORMATICS, 2018, 77 : 679 - 689
  • [39] AN EMPIRICAL EVALUATION OF DIMENSIONALITY REDUCTION USING LATENT SEMANTIC ANALYSIS ON HINDI TEXT
    Krishnamurthi, Karthik
    Sudi, Ravi Kumar
    Panuganti, Vijayapal Reddy
    Bulusu, Vishnu Vardhan
    2013 INTERNATIONAL CONFERENCE ON ASIAN LANGUAGE PROCESSING (IALP 2013), 2013, : 21 - 24
  • [40] Effect of dimensionality reduction on stock selection with cluster analysis in different market situations
    Han, Jingti
    Ge, Zhipeng
    EXPERT SYSTEMS WITH APPLICATIONS, 2020, 147