Reduced multidimensional scaling

被引:0
作者
Emmanuel Paradis
机构
[1] University of Montpellier,ISEM, IRD, CNRS, EPHE
来源
Computational Statistics | 2022年 / 37卷
关键词
Dimension reduction; Distance data; HIV; Multidimensional scaling;
D O I
暂无
中图分类号
学科分类号
摘要
Dimension reduction is a common problem when analysing large data sets. The present paper proposes a method called reduced multidimensional scaling based on performing an initial standard multidimensional scaling on a reduced data set. This method faces the problem of finding a representative reduced sample. An algorithm is presented to perform this selection based on alternating sampling in outlier areas and observations in high density areas. A space is then constructed with the selected reduced sample by standard multidimentional scaling using pairwise distances. The observations not included in the reduced sample are then projected on the constructed space using Gower’s formula in order to obtain a final representation of the whole data set. The only requirement is the ability to compute distances among observations. A simulation study showed that the proposed algorithm results performs well to detect outliers. Evaluation of running times suggests that the proposed method could run in a few hours with data sets that would take more than one year to analyse with standard multidimensional scaling. An application is presented with a dataset of 9547 DNA sequences of human immunodeficiency viruses.
引用
收藏
页码:91 / 105
页数:14
相关论文
共 65 条
[1]  
Abraham G(2014)Fast principal component analysis of large-scale genome-wide data PLoS ONE 9 e93766-42
[2]  
Inouye M(2005)Augmented implicitly restarted Lanczos bidiagonalization methods SIAM J Sci Comput 27 19-44
[3]  
Baglama J(2019)Dimensionality reduction for visualizing single-cell data using UMAP Nat Biotechnol 37 38-1016
[4]  
Lothar R(2018)A fast likelihood solution to the genetic clustering problem Methods Ecol Evol 9 1006-24
[5]  
Becht E(2018)The idm package: incremental decomposition methods in R J Stat Softw Code Snippets 86 1-48
[6]  
McInnes L(2019)Randomized matrix decompositions using R J Stat Softw 89 1-585
[7]  
Healy J(2019)MASS-UMAP: fast and accurate analog ensemble search in weather radar archives Remote Sens 11 2922-288
[8]  
Dutertre CA(1968)Adding a point to vector diagrams in multivariate analysis Biometrika 55 582-27
[9]  
Kwok IWH(2011)Finding structure with randomness: probabilistic algorithms for constructing approximate matrix decompositions SIAM Rev 53 217-137
[10]  
Ng LG(1964)Multidimensional scaling by optimizing goodness of fit to a nonmetric hypothesis Psychometrika 29 1-386